Tech Lead (Site Reliability Engineering)
17 hours ago
Nottingham
• We are seeking an accomplished and forward-thinking technical leader to join our Risk Intelligence organisation as the Tech Lead - Site Reliability Engineering (SRE) for the Risk Screening Product Line, • Based in Nottingham, UK, this role will lead the reliability engineering function for critical Screening applications supporting the EMEA region, • The successful candidate will be accountable for ensuring the availability, scalability, performance, and security of business-critical platforms while driving operational resilience and service excellence across a complex, cloud-native technology landscape, • As a key technical leader, you will combine deep expertise in Site Reliability Engineering, cloud operations, and modern software delivery practices with a passion for building high-performing engineering teams, • You will lead reliability strategies, incident management, automation initiatives, and continuous improvement efforts, while partnering closely with product, engineering, and business stakeholders to align technology outcomes with organisational objectives, • This role offers an exciting opportunity to influence engineering culture, mentor and develop talent, and establish best-in-class operational practices within a growing and innovative UK-based technology organisation, • Lead 24x7 reliability operations for Screening services, ensuring uptime, performance, and compliance across WC1 and World Check verify Applications, • Define and implement SLOs, SLIs, and error budgets; maintain reliability scorecards and drive improvements through engineering backlogs, • Act as Major Incident Commander during critical outages; lead blameless post-incident reviews and ensure learnings are institutionalized, • Drive automation-first mindset across incident response, deployment, compliance, and observability, • Collaborate with product and engineering teams to embed non-functional requirements and secure-by-design principles into delivery pipelines, • Co-own cloud reliability roadmap with platform teams; standardize tooling for observability, ITSM, and incident communication, • Ensure DR readiness, runbook quality, and resilience patterns are consistently applied across services, • Lead and mentor a team of SRE engineers, fostering a culture of ownership, learning, and engineering excellence, • Drive career development, performance management, and technical capability growth across the team, • Collaborate with HR and Talent teams to build local hiring pipelines and support workforce planning, • Promote well-being and psychological safety within the team; ensure compliance with health and safety standards, • Represent Nottingham site in global SRE forums; contribute to offshore strategy and location planning, • Partner with vendors and staffing partners to manage workforce augmentation and ensure delivery quality, • Support BCP/DR planning and ensure site-level operational readiness for critical events Empathetic and inclusive leader, committed to team well-being and growth • Familiarity with observability platforms (Datadog, BigPanda, OpenTelemetry), • Excellent communication and stakeholder management skills; ability to influence across technical and business domains, • Hands-on experience with container orchestration (Kubernetes, Docker), • Experience working with identity platforms and/or fraud detection systems, • Strategic thinker with a bias for action and accountability, • Strong understanding of SRE principles (SLOs, SLIs, error budgets, incident response), • AWS Lambda, ECS, RDS, CloudWatch, • Calm and structured under pressure; able to lead teams through high-stakes incidents, • Strong analytical mindset with a focus on measurable outcomes and continuous improvement, • Committed to continuous learning and knowledge sharing, • Proficiency in CI/CD pipelines and infrastructure-as-code tools (Terraform, GitHub Actions, Jenkins), • Strong collaborator who builds trust across teams and geographies, • 8+ years in production operations, SRE, or DevOps roles, with at least 3+ years in a people management capacity, • Proven experience managing cloud-native services on Azure and AWS, including: Azure SQL, Cosmos DB, Application Gateway, Key Vault, Storage, DNS, Load Balancer, Virtual Machines, Azure Machine Learning, Sentinel, • Passionate about engineering excellence, automation, and reliability #J-18808-Ljbffr