Director of Platform & Reliability
hace 6 horas
London
At Tembo, we’re on a mission to make home happen. We help people buy their first home, remortgage, move up the property ladder, and save towards homeownership through products including our Lifetime ISA, Cash ISA, and mortgage platform. Homeownership has become increasingly difficult for a generation of people. Through innovative mortgage products, affordability schemes, and savings tools, we help families who might otherwise be locked out of the property market find a way forward. Our work has a direct and visible impact on customers’ lives. We’re rated 4.8 on Trustpilot, with thousands of reviews from people we’ve helped get onto the property ladder or build a stronger financial future. Tembo is entering its next stage of growth. As our customer base, product offering, regulatory responsibilities, and engineering organisation expand, the reliability and resilience of our platform will become increasingly critical. We’re now looking for a Director of Platform & Reliability to take ownership of the technical foundations on which Tembo operates. This is a director-level individual contributor role reporting directly to the CTO. You will have company-wide responsibility for Tembo’s platform infrastructure, reliability, operational resilience, cloud architecture, and infrastructure engineering strategy. You’ll work closely with the CTO, Head of Engineering, and CISO to ensure our technology platform is secure, resilient, observable, cost-effective, and capable of supporting Tembo’s continued growth. This is not a people-management role, nor is it an architecture-only position. You will remain deeply hands-on: designing infrastructure, writing automation, improving production systems, responding to incidents, and delivering major platform initiatives yourself. At the same time, you will act as Tembo’s senior technical authority for platform and reliability. You’ll set direction, make important architectural decisions, establish standards, and help engineering teams build and operate reliable systems. The role goes significantly beyond traditional DevOps. You will be responsible for whether Tembo’s platform can withstand failures, recover from disruption, meet its regulatory obligations, and continue operating as the business scales. Define and execute Tembo’s platform and infrastructure strategy across our AWS and Azure environments. You’ll establish a clear target architecture and create a pragmatic roadmap for moving towards it without slowing down product delivery. Strengthen Tembo’s operational resilience. This will include: • Identifying critical services and technical dependencies, • Defining appropriate recovery objectives, • Improving disaster-recovery and business-continuity capabilities, • Designing and running resilience and recovery exercises, • Ensuring incidents, failures, and third-party outages can be handled effectively, • Producing the technical evidence needed to demonstrate that our controls work in practice Design, build, and operate the infrastructure that powers Tembo’s mortgage, savings, and internal operational platforms. You’ll take direct ownership of major infrastructure projects, including cloud migrations, platform modernisation, security improvements, and reliability initiatives. Improve how we operate production systems across the organisation. You’ll help establish clear approaches to: • Service ownership, • Monitoring and alerting, • Service-level objectives, • On-call practices, • Capacity and performance management, • Reliability engineering Drive the adoption and quality of infrastructure as code, primarily using Terraform. You’ll reduce manual configuration, improve consistency between environments, and ensure infrastructure changes are testable, reviewable, and auditable. Make it easier and safer for engineers to build, deploy, and operate software. You’ll work with engineering teams to improve CI/CD pipelines, deployment workflows, development environments, observability, and the internal tooling used to manage services. The goal is not to become a gatekeeper. It is to create strong foundations and paved paths that allow engineers to move quickly with confidence. Partner closely with the CISO to translate security requirements into practical technical controls. You’ll help design secure cloud architectures, reduce infrastructure risk, improve access and secrets management, and ensure security is built into our platform rather than added after the fact. Take a proactive approach to cloud cost, capacity planning, and vendor management. You’ll ensure Tembo understands where infrastructure spend is going and make informed decisions that balance cost, performance, resilience, and engineering productivity. • Establishing a long-term platform strategy across AWS and inherited Azure environments, • Preparing Tembo’s technical platform for increased regulatory and operational-resilience requirements, • Identifying and reducing single points of failure across critical customer journeys, • Defining recovery objectives and proving that critical services can recover within them, • Improving disaster-recovery planning and running meaningful resilience exercises, • Modernising legacy infrastructure supporting .NET and Ruby applications, • Expanding Terraform and infrastructure-as-code adoption, • Improving observability across mortgage, savings, and operational systems, • Creating clearer service ownership and incident-response processes, • Making deployments faster, safer, and easier for product engineers, • Improving security controls without creating unnecessary delivery friction, • Reducing cloud cost while improving reliability and performance, • Managing the technical risks created by rapid growth, third-party integrations, and increasingly complex financial products Within your first 6–12 months, you will likely have: • Established a clear platform and reliability strategy for Tembo, • Built a prioritised roadmap covering infrastructure, resilience, security, observability, and developer experience, • Taken ownership of Tembo’s operational-resilience engineering programme, • Identified critical systems, dependencies, failure modes, and recovery requirements, • Delivered and tested meaningful improvements to disaster recovery and service continuity, • Improved observability and production ownership across critical services, • Increased the consistency and coverage of infrastructure as code, • Reduced risk within inherited and legacy infrastructure, • Improved the reliability and security of our AWS and Azure environments, • Made it easier for engineering teams to deploy and operate their services, • Established yourself as a trusted technical partner to the CTO, Head of Engineering, CISO, and wider leadership team, • Taken direct ownership of some of Tembo’s most important and difficult technical initiatives We’re looking for an experienced platform, infrastructure, or reliability engineer who has operated at principal, staff, head-of, or director level while remaining deeply hands-on. You will likely have: • Significant experience designing and operating production infrastructure, • A track record of owning platform or reliability strategy across an organisation, • Deep experience with AWS, Azure, or complex multi-cloud environments, • Strong experience with Terraform and infrastructure-as-code practices, • Experience designing and testing disaster-recovery and operational-resilience capabilities, • Strong knowledge of observability, incident response, service reliability, and production operations, • Experience leading major infrastructure migrations or modernisation programmes, • A strong understanding of cloud security and infrastructure risk, • Experience improving CI/CD and developer-platform capabilities, • The ability to move between long-term strategy and detailed implementation, • Strong judgement about where standardisation is valuable and where teams need flexibility, • The communication skills to work effectively with engineering, security, product, operations, risk, and senior leadership, • The confidence to challenge decisions constructively and take responsibility for the outcome, • A willingness to remain hands-on and personally deliver critical technical work We’re particularly interested in people who have taken genuine accountability for the resilience and operation of a platform, rather than only advising teams or managing infrastructure backlogs. Your work at Tembo will directly affect our ability to help people get onto the property ladder and build a stronger financial future. You’ll have significant influence over the technical foundations of the company and work closely with the CTO, Head of Engineering, CISO, and wider leadership team. You’ll be joining at an important point in Tembo’s growth, with the opportunity to shape how we approach platform engineering, security, reliability, and operational resilience for years to come. This is a rare opportunity to take on director-level scope while remaining an individual contributor and staying deeply involved in the technology. • Private health insurance, • Enhanced parental benefits, • Birthday day off, • Annual international team offsite, • Competitive salary and equity We innovate relentlessly and leave no stone unturned in helping customers find a way onto the property ladder. Customers, colleagues, and partners are at the heart of everything we do. We listen, challenge assumptions, and continuously improve how we work. #J-18808-Ljbffr