Service Manager – Site Reliability Engineering
20 hours ago
Belfast
• Own the overall health, reliability, availability, and observability of critical business applications and technology services, • Establish, monitor, and report on SLAs, SLOs, Error Budgets, availability, performance, and operational KPIs, • Lead Major Incident Management activities, coordinating cross-functional teams during outages and ensuring rapid service restoration and RCA completion, • Drive Problem Management by identifying recurring issues, analyzing systemic failures, implementing corrective actions, and reducing operational risk, • Govern Change and Release Management processes, production readiness reviews, maintenance, and deployment activities, • Ensure effective observability through monitoring, alerting, logging, dashboards, and operational reporting, • Promote automation and operational efficiency through Infrastructure as Code, DevOps, self-healing, and auto-remediation, • Partner with engineering, platform, infrastructure, security, and business teams to improve resilience, scalability, stability, and customer experience, • Act as operational liaison for business stakeholders and vendors, conducting service reviews, communicating risks, and managing escalations, • Ensure adherence to ITIL-based processes, governance standards, compliance requirements, audit obligations, and operational documentation standards, • Lead continuous service improvement initiatives to reduce MTTR, increase stability, improve customer satisfaction, and enhance operational maturity, • Lead operational governance ceremonies, service reviews, incident and problem reviews, change governance forums, readiness assessments, stakeholder communications, and executive service reporting, • Create, maintain, review, and ensure compliance with runbooks, SOPs, knowledge articles, audit evidence, disaster recovery procedures, and certification artifacts, • Provide leadership across incident, problem, change, release, and service management disciplines, • Lead, mentor, coach, and develop a team of Service Analysts, • Provide training, knowledge sharing, cross-training, and professional development support, • Support operational risk management by identifying vulnerabilities, assessing impacts, and developing mitigation plans, • Collaborate on service reliability, scalability, security, production readiness, modernization, and transformation initiatives, • Drive strategic operational improvements supporting service quality, customer experience, business alignment, and long-term sustainability, • Legal right to work in the UK; Allstate is not providing sponsorship for this vacancy, • Minimum of 4 years of experience supporting or improving enterprise technology services, infrastructure environments, platform operations, Site Reliability Engineering, IT Operations, or Service Management disciplines (or equivalent), • Minimum of 2 years leading and mentoring teams within service reliability, availability, performance, and/or operational governance in a large enterprise environment, • Experience leading Major Incident response activities, coordinating cross-functional teams, and driving RCA efforts and corrective actions, • Experience implementing operational improvements, automation initiatives, risk-reduction measures, and continuous service improvement programs, • Strong understanding of observability practices, including monitoring, alerting, logging, dashboards, and operational reporting, • Knowledge of ITIL principles and IT Service Management processes, • Experience developing and maintaining operational documentation, runbooks, support procedures, recovery documentation, and knowledge articles, • Experience working within an SRE, DevOps, Cloud Operations, Platform Engineering, Enterprise Operations, or production support environment, • Experience supporting Identity and Access Management platforms, including IAM, ISAM, IBM Verify, SailPoint, or related identity technologies, • Experience managing SLAs and operational KPIs (or equivalent), • Knowledge of networking technologies, firewalls, DNS, load balancing, and enterprise infrastructure concepts, • Experience leading or mentoring a team of engineers (or equivalent), • Experience with Infrastructure as Code, automation frameworks, cloud-native operational practices, self-healing systems, or auto-remediation, • Familiarity with API management platforms and enterprise service-integration technologies, • Experience supporting messaging or event-streaming platforms such as Kafka, • Knowledge of middleware or integration technologies such as TIBCO or comparable enterprise platforms, • Experience supporting Microsoft Azure, Amazon Web Services, or Google Cloud Platform, • Experience using ServiceNow or a comparable ITSM platform, • Knowledge of compliance, risk management, audit controls, operational resilience, business continuity, and disaster recovery practices, • Relevant professional certifications, such as ITIL, SRE, cloud, Kubernetes, security, ServiceNow, or comparable technology certifications, • Experience leading operational maturity assessments, service governance programs, or production readiness reviews Demonstrates expertise in Service Reliability Engineering, Incident Management, and ITIL-based processes, with a strong focus on operational governance, automation, and continuous service improvement. Proven ability to lead cross-functional teams, manage SLAs, and enhance customer experience through effective communication and collaboration. • Service Reliability Engineering, • Incident Management, • ITIL Principles, • Infrastructure as Code, • Operational Governance, • Operational Documentation, • Monitoring, • Alerting, • Logging, • Dashboards, • Cloud Operations, • Automation Frameworks, • Identity and Access Management, • Networking Technologies, • Event-Streaming Platforms, • Leadership, • Mentoring, • Collaboration, • Communication, • Problem-Solving, • ITIL, • SRE, • Cloud Certifications, • Kubernetes, • Security Certifications, • Operational Resilience, • Business Continuity, • Disaster Recovery, • Compliance, • Risk Management, • ServiceNow, • Microsoft Azure, • Amazon Web Services, • Google Cloud Platform, • Kafka, • TIBCO #J-18808-Ljbffr