Senior Site Reliability Engineer
6 days ago
Sant Cugat del Vallès
ph3Position Overview /h3pWe are building a global Site Reliability Engineering (SRE) team to support critical commercial and internal platforms and applications. As an SRE, you will design, build, and scale reliable distributed systems that power healthcare innovation worldwide. The role focuses on reliability, scalability, automation, and operational excellence and includes participation in a structured on‑call rotation. /ph3Core Responsibilities /h3ulliDefine and implement SLIs, SLOs, and error budgets with product and engineering teams. /liliConduct reliability reviews for new and existing services. /liliDesign scalable, fault‑tolerant architectures in AWS and Azure environments. /liliLead capacity planning, performance and cost optimization initiatives. /liliImprove system resilience through automation and self‑healing patterns. /liliDrive organizational observability maturity (metrics, logs, traces, alert quality). /liliPerform complex root‑cause analysis and drive rapid mitigation. /liliParticipate in blameless post‑mortems and follow‑through. /liliImprove MTTR, reduce incident frequency, and elevate production standards. /liliCollaborate seamlessly with engineering teams to enable timely and effective resolutions. /liliHandle requests and incidents, create and maintain runbooks. /liliParticipate in a structured 24/7 on‑call rotation. /liliReduce operational toil through tooling and automation (Python or similar). /liliImprove CI/CD reliability and deployment safety mechanisms. /liliBuild and maintain infrastructure‑as‑code (Terraform or equivalent). /liliEnhance Kubernetes platform reliability (EKS, AKS, or similar). /liliPartner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycle. /liliMentor mid‑level engineers and help shape SRE best practices. /liliChampion a culture of ownership, accountability, and continuous improvement. /li /ulh3Qualifications /h3ulliBachelor’s degree in computer science, engineering, or a related field, or equivalent professional experience. /liliExperience in site reliability engineering, software engineering, or related fields with production on‑call experience. /liliSolid experience with AWS and/or Azure, including setting up, monitoring, and maintaining cloud resources (incl. Kubernetes, EKS, AKS, GKE). /liliProficiency with observability tools. /liliHands‑on experience with incident management tools. /liliProficiency in scripting languages for automation purposes (Python, etc.). /liliDemonstrated proficiency in troubleshooting, especially in cloud and distributed system environments. /liliExcellent communication, teamwork and documentation skills, with a proactive and self‑motivated approach to improving system reliability and operational efficiencies. /liliProficient spoken and written English communication. /li /ulh3Location /h3pPrimary location: Sant Cugat del Vallès. Additional locations may be available. /ph3Equal Opportunity Employer /h3pRoche is an Equal Opportunity Employer. We believe it’s urgent to deliver medical solutions right now – even as we develop innovations for the future. We are passionate about transforming patients’ lives. We are courageous in both decision and action. And we believe that good business means a better world. We are committed to scientific rigor, unassailable ethics, and access to medical innovations for all. We are proud of who we are, what we do, and how we do it. We are many, working as one across functions, across companies, and across the world. /p /p #J-18808-Ljbffr