Site Reliability Engineer
hace 2 días
Madrid
At Tinybird, we help developers and data teams unlock the power of real‑time data. We enable the rapid building of data pipelines and innovative data products by ingesting multiple data sources at scale, querying them with familiar SQL, and publishing low‑latency, high‑concurrency APIs for applications. Developers can create fast APIs in minutes, transforming what once took hours or days. The Platform team builds, operates, and continually improves the technical foundations that Tinybird relies on. We ensure safety at scale, predictable evolution, and product and customer growth with minimal operational friction. Our focus is on reliability, observability, performance, infrastructure, cost efficiency, CI/CD, development environments, and critical backend services. We are both an infrastructure and operations team, making Tinybird’s foundations reliable, observable, scalable, and easier to evolve. Role: Site Reliability Engineer who loves keeping large‑scale distributed systems reliable and adaptable as they grow. You should understand hardware and software, grasp our product, and tackle real challenges faced by customers and internal teams. • Strong experience designing, building, and running distributed cloud architectures and large‑scale web‑based production systems., • Deep knowledge of Kubernetes – essential. Comfortable designing and operating production‑grade clusters, writing custom controllers or operators, and tuning autoscaling mechanisms (e.g., KEDA, Karpenter)., • Knowledge of how Kubernetes manages networking, storage, scheduling, and resources; able to analyze performance and failure scenarios at scale., • Skilled in AWS and GCP., • Baseline coding skills (Python or C++). Able to explore our codebase, ClickHouse source code, and understand how things work., • Experience operating close to production: debugging incidents, analyzing system behavior, improving observability, and enhancing service reliability., • Systems thinking with attention to edge cases, failure modes, and implementation details. Careful about performance, reliability, cost efficiency, and operational simplicity., • Action orientation – quick decisions, iterative delivery, and ownership over issues that need fixing., • Curiosity about data and SQL; comfortable querying our own data. Experience with ClickHouse or launching databases at scale is a plus., • Familiarity with Traefik, Varnish, Redis, Terraform, or Ansible is helpful but not mandatory., • Strong written communication for asynchronous work and documentation., • Fluency in English and Spanish (English is primary; Spanish is common in the Platform team)., • Willingness to participate in on‑call rotations and located in an EU timezone., • Kubernetes – foundation of our infrastructure., • ClickHouse – primary data store., • Python – backend; some performance‑critical components in C++., • Varnish – load balancing and caching., • Traefik – ingress and routing., • Redis – metadata store., • Zookeeper – coordinating ClickHouse replicas., • ArgoCD – GitOps‑based continuous delivery., • Grafana, Loki, Mimir, OpenTelemetry – monitoring, alerting, telemetry., • Enhance high‑availability and elasticity for automatic, efficient scaling., • Improve observability across resource usage, service metrics, telemetry, dashboards, and alerts., • Strengthen disaster recovery tools, incident discovery, and on‑call experience., • Handle Kubernetes lifecycle tasks, cluster infrastructure, autoscaling, and safe deployments., • Understand ClickHouse internals and extract peak performance., • Identify bottlenecks and improve performance across storage, networking, and compute., • Reduce operational burden by automating and managing processes., • Assist with incident prevention, post‑reliability reviews, and follow‑up tasks., • Strengthen CI/CD foundations for greater confidence in deployments., • Collaborate with product and backend teams to design architecture, optimize resources, and increase platform autonomy. Your focus remains on the Platform, but priorities often come from product, customers, and engineering teams. You might design autoscaling behavior, develop Kubernetes infrastructure, investigate production issues, improve dashboards, enable safer deployments, optimize ClickHouse, or help reduce platform friction for other teams. We are a remote‑first company with occasional in‑person gatherings in Madrid to foster alignment and problem solving. We value ownership, transparency, clear communication, and documented decisions. Work is closely aligned with product, support, customer success, and other engineering teams to support the entire company. • 22 days of holiday a year plus birthday and public holidays., • Freedom to work from wherever suits you best., • Up to €2,800 for home‑workspace setup., • Remote‑first culture with flexibility. #J-18808-Ljbffr