
Scaleway · Paris
OUR STORY: 🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supp...
🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow !
Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.
Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.
With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.
Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.
📍 Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.
Our growth is driving us to strengthen our Network SRE Products team to ensure the high reliability, performance, and scalability of our storage platforms.
Your mission will be to automate, monitor, and improve the reliability, performance, and scalability of our infrastructure. You will maximize availability, optimize fault tolerance, and reduce operational overhead—ensuring robust and efficient systems for our products and services.
We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together. You will be part of a team of Site Reliability Engineers reporting to a Lead SRE and integrated into the SRE Guild, a collective focused on fostering best practices across engineering.
The team collaborates daily with Dev, Product, and Ops teams to improve resiliency, support service scalability, and ensure a seamless customer experience across our network solutions.
Hybrid work: We offer up to 3 days of remote work per week.
Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
Initial call with a recruiter to get to know each other (30 min)
Technical interview with the Head of SRE to assess your skills and approach (1h)
Manager x team interview to explore your background and team fit (1h)
Final interview with HR and office visit to meet your future teammates and discover our workspaces
Version française disponible ici
WHY THIS ROLE As our Director of Site Reliability Engineering, reporting to our VP of Platform Engineering, you'll own the core infrastructure layers that everything at Doctolib runs on: cloud infrastructure, database operations, network infrastructure, and observability. You will also lead the Doctolib Operations Center (DOC) and drive a decisive shift from reactive operations to a proactive, world-class reliability culture. This is a rare opportunity to shape the infrastructure backbone of Europe's leading healthtech company, at a moment when Doctolib is actively expanding multi-cloud capabilities, scaling to new countries, and building the reliability culture that will define the next decade of healthcare innovation. Why this is an extraordinary challenge * Real stakes, every day. When Doctolib is down, consultations don't happen, diagnoses are delayed, care journeys are interrupted. The infrastructure you build is a direct lever on patient outcomes — in a world where 8 of the top 10 causes of death in Europe are preventable. * A once-in-a-generation platform transition. Multi-cloud, monolith modularisation, international expansion — all happening simultaneously. You won't inherit a finished platform. You'll define what it becomes. * Reliability as the competitive moat. As we scale AI health companions, automate clinical workflows, and launch across Europe, the speed and resilience of the platform directly determines how fast 700+ engineers can ship innovations that change healthcare. * A cultural build, not just a technical one. The incident response culture, observability standards, and operational ownership model you establish here will shape how Doctolib engineers work for years to come. WHAT YOU'LL DO * Build and run a world-class SRE org of 25+ engineers across Cloud Infrastructure, Database & Storage, Network Infrastructure, Observability Tooling, and the Doctolib Operations Center * Own the infrastructure strategy and roadmap — cloud, database, network, observability — and deliver against company OKRs * Lead the Doctolib Operations Center: set incident response standards, drive MTTR reduction, embed blameless post-mortem culture across engineering * Architect and execute our multi-cloud strategy — reducing vendor lock-in, cutting migration costs, and enabling international expansion * Own network infrastructure at scale: load balancing, CDN/WAF, VPCs, peering, zero-trust networking across a high-traffic, multi-country platform * Drive observability as a product — give 700+ engineers true visibility into system health and turn observability maturity into an operational excellence lever * Lead from the front as a senior technical voice in the Platform org and broader Tech leadership team WHO YOU ARE * 12+ years in software engineering, including 5+ years leading managers and running infrastructure or SRE organisations at scale * Track record of taking SRE practices from reactive to proactive — with measurable reductions in incidents and MTTR * Strong multi-cloud and network infrastructure experience: load balancing, CDN/WAF, VPCs, peering, at high-traffic scale * Deep database operations background: large-scale transactional systems (PostgreSQL, Aurora), streaming/CDC (Kafka), data layer FinOps * Experience building observability platforms that give teams genuine visibility — metrics, logs, traces, alerting * Sharp process thinking: SLOs, error budgets, incident management, blameless post-mortems * Outcome-driven: you track reliability, cost efficiency, and engineering velocity as business metrics, not just technical ones * Strong communicator and influencer at executive level — equally credible with senior engineers and business stakeholders * Builder of high-performing, people-first engineering cultures * Fluent in English; comfortable in fast-paced, international environments * You recognise yourself in our playbook values Bonus Points If You Have… * Experience in healthcare, regulated, or high-compliance industries (HDS, ISO 27001, SOC2, GDPR, data sovereignty) * Familiarity with our stack: Ruby on Rails, Node.js, Go, Python, React, AWS, GCP, Kubernetes, PostgreSQL, Datadog, GitHub Actions * French language proficiency * Experience with AI-augmented infrastructure tooling or ML platform operations * M&A or post-acquisition infrastructure integration experience WHAT WE OFFER * Free comprehensive health insurance for you and your children * Parent Care Program: receive additional leave on top of the legal parental leave * Free mental health and coaching services through our partner Moka.care * For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support * Work from abroad for up to 10 days per year thanks to our flexibility days policy * Work Council subsidy to refund part of sport club membership or creative class * Up to 14 days of RTT * A subsidy from the work council to refund part of the membership to a sport club or a creative class * Lunch voucher with Swile card If you would like to find out more about tech life at Doctolib, feel free to read our latest Medium blog articles! At Doctolib, we are committed to improving access to healthcare for everyone. This translates into our recruitment process. We evaluate candidates based solely on qualifications and motivation, without any form of discrimination. The more diverse ideas are heard, the more our product will truly improve healthcare for all. You are welcome to apply to Doctolib, regardless of your gender, religion, age, sexual orientation, ethnicity, disability. To ensure equal opportunities, we invite you to exclude personal information (e.g. pictures, age) from your applications. If you require any accommodation, please let us know for support during the hiring process. Join us in building the healthcare we all dream of! All information provided is processed by Doctolib for application management. For data processing details, click here. Please contact hr.dataprivacy(at)doctolib.com for inquiries or to exercise your rights. #LI-DB1
OUR STORY: 🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies. Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector. With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants. Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen. 📍 Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon. WHY WE NEED YOU? Our growth is driving us to strengthen our SRE team to support and scale our production environments. Your mission will be to build and maintain reliable, observable, and secure infrastructure in order to ensure optimal service availability for our customers around the world. YOUR FUTURE TEAM We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together. You will be part of a team of experienced Site Reliability Engineers. The team is responsible for maintaining and evolving core infrastructure and observability tools, supporting product teams, and improving the reliability of Scaleway’s services. YOUR DAILY ROUTINE - Build and optimize tooling to automate monitoring, diagnosis, and remediation of production incidents - Troubleshoot high-impact production issues in collaboration with other engineering teams - Participate in an on-call rotation to handle incidents and ensure service continuity - Implement and maintain observability solutions to monitor infrastructure and application health - Contribute to infrastructure lifecycle management across different environments - Promote and apply best practices in terms of stability, resiliency, scalability, and security - Maintain clear technical documentation for tools and procedures - Contribute to system and tool evolution based on production feedback - Collaborate closely with development teams to ensure infrastructure readiness - Participate in team rituals and knowledge-sharing initiatives ABOUT YOU SOFTSKILLS : - Proactive and solution-oriented mindset - Passion for automation and continuous improvement - Strong collaboration and communication skills - Ability to work independently and in a team - Willingness to mentor and share knowledge 💻 HARDSKILLS : - Experience with Go, Python or Rust - Strong scripting skills (Bash, Python) - Hands-on experience with Linux systems (Ubuntu/Debian) - Knowledge of networking (TCP/IP, DNS, BGP, load-balancing, IPv6, etc.) - Experience in cloud environments and infrastructure (bare metal, VMs, containers, orchestrators) - Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.) - Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.) - Experience managing relational databases (PostgreSQL) - Understanding of CI/CD pipelines (GitLab) - Comfortable with English (written and spoken) WHAT YOU WILL FIND AT SCALEWAY ++++ - Hybrid work: We offer up to 3 days of remote work per week. - Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities. - Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches. - Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life. International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French. - Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers. 🚀 Why join the Scaleway adventure ? ✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI. ✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges. ✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification. 🔜 THE NEXT STEPS … - Discovery call with a recruiter (30 min) - Interview with the manager to understand your technical skills and approach to the role (45 min) - Technical interview to validate your expertise (1h) - Interview with the Head of the Tribe to deepen your discussions and assess your fit with the team (45 min) - HR interview to tour our offices and meet your future colleagues
YOUR IMPACT We are looking for a Senior Site Reliability Engineer to join the Core Reliability & Observability team in Platform Engineering. Your mission will be to shape Doctolib's observability strategy and ensure our platform remains reliable, debuggable, and scalable at a European scale. You will work in a feature team developing logging, metrics, tracing, and alerting capabilities, contributing directly to supporting 400,000 health professionals and 80 million patients in their daily healthcare journey. Working in the tech team at Doctolib means building innovative products and features to improve the daily lives of care teams and patients. WHAT YOU'LL DO Your responsibilities include but are not limited to: * Lead the observability strategy across the platform, with an emphasis on building scalable, developer-friendly logging and tracing capabilities * Identify and lead large-scale cross-cutting reliability initiatives, including improvements to our incident detection, response, and postmortem analysis capabilities * Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry WHO YOU ARE Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply. You'll be a great fit if you: * Have a solid hands-on experience (3y+) on a large-scale production platform * Have proven experience with cloud platforms such as AWS, Azure or Google Cloud * Have solid understanding of containerization and orchestration technologies (Docker and Kubernetes) * Have a strong understanding of Helm for managing Kubernetes manifests and ArgoCD for GitOps workflows * Have deep expertise in observability tooling and architecture, such as: * Logging: Fluent Bit, OpenTelemetry, Loki, Elasticsearch, Logstash, Vector * Tracing: OpenTelemetry or proprietary APMs * Metrics: Prometheus, Thanos, Datadog, or equivalent * Have proficiency in at least one programming language (Ruby, Python, Go, Java, etc.) and a deep understanding of infrastructure as code principles * Have experience with monitoring and observability tools * Like troubleshooting performance issues in complex environments * Are fluent in English It would be fantastic if you: * Have experience contributing to open-source observability projects * Have worked in a high-growth tech environment * Are passionate about developer experience and platform engineering LIFE AT DOCTOLIB TECH * Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to country and healthcare specialty requirements. * Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native. * We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision here. Want to learn more about our tech culture and environment? Visit the Doctolib Tech site. WHAT WE OFFER * Free comprehensive health insurance (basic package) for you and your children * 25 days of paid vacation per year, plus up to 14 days of RTT * Free mental health and coaching services through our partner Moka.care * Work from abroad for up to 10 days per year thanks to our flexibility days policy * Lunch vouchers (Swile card) worth €8.50 per working day, with €4.50 covered by Doctolib * A subsidy from the work council to refund part of the membership to a sport club or a creative class * 50% reimbursement of your public transport subscription * ParentCare Program: Enjoy full salary coverage (100%) during your first month of birth leave, and 75% during the second, covered by Doctolib * Enrollment in Doctolib's long-term employee value sharing plan called DoctoGrowth * For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support * Relocation support in case of international mobility * Access to the best AI tools for coding, development and dedicated training OUR INTERVIEW PROCESS * Recruiter Interview * Technical SRE Interview * System Design Interview * Behavioral Interview * At least one reference check We want your experience to be clear, respectful, and transparent. Learn more about our hiring process on our candidate experience page. JOB DETAILS * Permanent position * Tech stack: Kubernetes, Prometheus, OpenTelemetry, Loki, ArgoCD, Ruby, Python, Go * Full-time * Paris, France * Hybrid work setup (up to 2 remote days per week) * Start date: as soon as possible WE WELCOME EVERYONE At Doctolib, we are committed to improving access to healthcare for everyone. This translates into our recruitment process. We evaluate candidates based solely on qualifications and motivation, without any form of discrimination. The more diverse ideas are heard, the more our product will truly improve healthcare for all. You are welcome to apply to Doctolib, regardless of your gender, religion, age, sexual orientation, ethnicity, or disability. To ensure equal opportunities, we invite you to exclude personal information (e.g., pictures, age) from your applications. If you require any accommodation, please let us know for support during the hiring process. Join us in building the healthcare we all dream of! YOUR DATA PRIVACY All information provided is processed by Doctolib for application management. For data processing details, click here: Germany l France l Italy l Netherlands. Please contact hr.dataprivacy(at)doctolib.com for inquiries or to exercise your rights.