Senior SRE Engineer
hace 1 día
Barcelona
ph3About Codeway /h3 pCodeway is a global consumer tech company with more than 400 million users worldwide. Since 2020, we’ve built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning, and entertainment. Our flagship apps — Retake AI, Cleanup, Learna, and DramaPops — lead their categories globally. In 2024, we became the most downloaded app publisher on iOS. We’re a team of 300+ people across İstanbul and Barcelona who bring curiosity, passion, trust, and ownership to everything we build. Recognized as a #1 LinkedIn Top Startup and a Great Place to Work in Europe, Codeway is where ambitious people do their life’s best work. We’re building the next generation of consumer tech and reimagining what mobile apps can be. This is Codeway. This is our way. Join us. /p h3Position /h3 pWe’re looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world. /p h3What You’ll Be Doing /h3 h3Reliability, SLOs, Observability /h3 ul liDefine, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion. /li liOwn observability end‑to‑end — metrics, logs, traces, dashboards, and alerting — and drive measurable reductions in detection and resolution times. /li liReduce alert noise and false positives so on‑call engineers can trust what wakes them up. /li liRun reliability reviews and an error‑budget policy that shapes how teams prioritize between shipping and stability. /li /ul h3Kubernetes Platform Operations /h3 ul liOperate, scale, and upgrade our multi‑cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management. /li liAct as the deep‑expertise escalation point for cluster and platform issues across dozens of services. /li liOwn capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets. /li liBuild self‑service platform tooling that lets product teams move quickly without needing to become infrastructure experts. /li /ul h3Security Resilience /h3 ul liEmbed security into the platform through RBAC and least‑privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence. /li liPartner with the security function on vulnerability remediation, audit readiness, and secure‑by‑default infrastructure. /li liOwn disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing. /li liContribute to architecture and production‑readiness reviews so reliability and security are designed in, not bolted on. /li /ul h3Incident Response Automation /h3 ul liLead the on‑call rotation and act as incident commander during production incidents. /li liRun blameless post‑mortems, quantify impact, and track corrective actions through to closure so the same failure doesn’t recur. /li liBuild and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks. /li liSystematically identify, measure, and eliminate operational toil through automation, protecting engineering time for high‑leverage work. /li /ul h3WHAT YOU’LL BRING? /h3 ul liExperience operating high‑traffic, always‑on production systems at meaningful scale, typically gained over 5–8 years in SRE, Platform, or DevOps roles. /li liHands‑on production Kubernetes experience — you’ve run clusters day to day, through upgrades, autoscaling, and real troubleshooting under load, not just deployed to them. /li liA strong cloud engineering background, along with solid Linux and networking fundamentals. /li liA track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes. /li liExperience with Infrastructure as Code and CI/CD pipeline design — you treat infrastructure and delivery as code. /li liDepth in observability tooling: instrumentation, dashboarding, and alert design. /li liA genuine security‑first mindset, where least‑privilege, secrets hygiene, and vulnerability management are habits rather than afterthoughts. /li liScripting and automation fluency in at least one language, used to build tooling and remove toil. /li liIncident‑command experience: owning on‑call, running blameless post‑mortems, and driving resolution times down over time. /li liAbility to communicate clearly with both engineers and leadership, especially under pressure. /li /ul h3NICE TO HAVE /h3 ul liExperience with high‑scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics. /li liExperience with GitOps and progressive‑delivery patterns such as canary and blue‑green rollouts. /li liFamiliarity with service mesh, API gateways, or multi‑region and multi‑cluster topologies. /li liCloud cost optimization and FinOps discipline at scale. /li liExposure to compliance initiatives (SOC 2, ISO 27001, GDPR) and broader DevSecOps practice. /li liChaos engineering or resilience testing experience. /li liRelevant certifications in Kubernetes, cloud, or DevOps disciplines. /li liExperience supporting many independent services and teams concurrently in a fast‑shipping, product‑led environment. /li /ul h3OUR ENVIRONMENT /h3 ul liKubernetes (GKE) and containerized workloads /li liGoogle Cloud Platform (GCP) /li liTerraform and Infrastructure as Code /li liCI/CD and GitOps tooling /li liModern observability (metrics, logs, traces, alerting) /li liCloudflare CDN and edge /li /ul h3What Success Looks Like /h3 ul liClear SLIs, SLOs, and error budgets live and reported for our most critical services. /li liA measurable reduction in detection and resolution times for production incidents. /li liA consistent, blameless incident response practice, with postmortem actions tracked to closure. /li liA hardened Kubernetes fleet, with key security gaps closed across access control, secrets, scanning, and patching. /li liExpanded Infrastructure as Code and observability coverage across the platform. /li liDisaster recovery drills that reliably pass agreed RTO/RPO targets. /li liOperational toil trending down against an explicit target, with automation replacing manual work. /li liProduct teams self‑serving on standardized, secure, well‑instrumented platform tooling. /li liShared measurement with the Security team, so security outcomes are tracked with the same rigor as reliability ones: vulnerability remediation times, patching cadence and coverage, access review completion, and security incident detection and response times — reported as real metrics, not status updates. /li liA single, agreed view of operational health across engineering and security, so the two functions work from the same definitions and the same dashboards rather than competing narratives. /li liVisibility into how efficiently we operate as an organization, connecting platform performance to the business it supports — the reliability of the systems our marketing and growth stack depends on, the speed and safety of our path from commit to production, and the infrastructure cost behind serving and acquiring users. /li liMetrics leadership actually uses, giving the wider company a clear, honest picture of where we’re strong, where we’re exposed, and where the next investment should go. /li /ul h3What We Offer /h3 ul liA Competitive Compensation Package. Long story short, we take care of you. /li liA Meal Compensation that covers a decent and nutritious lunch. /li liFull health benefits, including unlimited private health insurance and HPV vaccine coverage. /li liPet adoption support covering primary healthcare, parasite vaccinations, and microchip costs in the first year after adoption. /li liState‑of‑the‑art tech: MacBook, iPhone 15 Pro, Magic Mouse, Magic Keyboard, adjustable desk, 4K screen, and any other gadget you need. /li liSport activities support and gym membership. /li liFlexible schedule that encourages responsibility over tracking every minute. /li liEnglish language course support. /li liTop‑notch office in the heart of Barcelona at the iconic Edifici Estel. /li liCodebrew coffee shop with healthy snacks at all hours and free breakfast lunch daily. /li liNo dress code. /li liA dynamic team environment with young, talented, and passionate members. /li liGaming area with PS5 corner. /li liSoftware support subscriptions. /li liPublic transportation support with additional monthly compensation. /li /ul h3Recruiting Process /h3 ul liApplication: Send us your CV or LinkedIn profile. You can also write a few words about yourself. /li liHiring Manager Interview. /li liTalent Culture Interview. /li liCase Study. /li liTechnical Interview. /li liWelcome aboard! You are now part of the team. /li /ul pYour personal data that will be collected through your applications made due to this advertisement will be processed automatically by the data controller Codeway in accordance with the Personal Data Protection Law No. 6698 (“KVKK”) and relevant legislation for the purpose of conducting job application processes. You can access detailed information on this matter from the Candidate Information Notice page at Codeway processes the personal data of job applicants in accordance with relevant legislation. Privacy Notice for Employee Candidate summarizes the data processing procedures during the application process. /p /p #J-18808-Ljbffr