Sovereign Engineering Platform SRE (m/f/d)
17 days ago
Valencia
T‑Systems is part of the Deutsche Telekom Group, with around 30.000 employees worldwide. We create technology with purpose to generate a positive impact on society. We are looking for curious talent, eager to learn, take on challenges, and contribute ideas that transform our customers’ experience. We trust people: we offer autonomy, continuous support, and a collaborative environment where you can grow without limits. We are one global team, guided by respect, integrity, and a passion for doing better every day. • Build and operate Kubernetes environments that host AI engineering tools, internal model gateways, retrieval components, workflow services, CI/CD runners, and documentation services., • Implement GitOps and Infrastructure as Code patterns for reproducible provisioning, configuration, policy enforcement, platform upgrades, and disaster recovery readiness., • Manage private registries, package mirrors, secrets, identity integration, network segmentation, storage classes, backup routines, and controlled connectivity models., • Provide observability for engineering workloads, including metrics, logs, traces, GPU and CPU utilization, service health, cost signals, and operational runbooks., • Work with software, security, and architecture teams to ensure the platform supports AI-assisted SDLC workflows without creating uncontrolled data exposure or audit gaps., • Platform tooling such as Kubernetes, Helm, Terraform, Ansible, ArgoCD, Crossplane, GitLab runners, Jenkins agents, private registries, and internal package mirrors., • AI platform components such as vLLM, Ollama, OpenAI-compatible gateways, Qdrant or similar vector stores, Open WebUI, Continue-compatible endpoints, and workflow services., • Observability and operations stacks such as Prometheus, Grafana, Loki, OpenTelemetry, ELK/OpenSearch, Alertmanager, SRE runbooks, and incident management tooling., • Security and governance components such as Vault, Keycloak, network policies, RBAC, admission controls, image scanning, SBOM tooling, and audit logging., • Infrastructure awareness covering GPU-backed nodes, CPU-only fallback, storage performance, network isolation, proxy patterns, on-premise environments, and dedicated landing zones., • 5+ years in SRE, platform engineering, DevOps, cloud infrastructure, or operations roles with strong Kubernetes and Linux expertise., • Proven experience building and operating production-grade engineering platforms with GitOps, Infrastructure as Code, observability, and operational runbooks., • Hands‑on skills in Terraform, Ansible, Helm, Python or shell scripting, CI/CD runners, private registries, and secure configuration management., • Good understanding of networking, storage, secrets, access control, monitoring, backup, disaster recovery, and operational hardening in high-security environments., • Comfortable supporting AI‑enabled engineering workloads in sovereignty‑driven contexts where isolation, controlled data handling, reliability, and auditability are mandatory., • International, dynamic and collaborative environment., • T‑Social: social initiatives (sports, community, health, ...)., • Hybrid work model (remote/on‑site)., • Flexible working hours., • Customized training: access to Coursera to learn whatever you want, whenever you want., • Weekly language classes (English & German)., • International Mentoring Sessions & Experience Days., • Flexible compensation plan (health insurance, meal vouchers, childcare, transport)., • Telemedicine., • Life and accident insurance., • Social fund., • 26+ working days of vacation per year., • Free access to specialist services (medical, legal, wellness)., • 100% salary coverage during medical leave. T-Systems Iberia will only process the CVs of candidates who meet the requirements specified for each offer. #J-18808-Ljbffr