Site Reliability Engineer (Guardicore AI Platform) - Remote
hace 10 horas
Córdoba
Overview In this role you will own the reliability and operational readiness of the Guardicore Data and AI Platform, a cloud-native data/AI platform for security analytics. You will work with cross-functional teams to improve availability, performance, security, and cost efficiency, while guiding engineers on service performance. You will lead complex production investigations and leverage AI-driven automation to streamline operations. This is an opportunity to shape platform reliability at scale within a security-focused, AI-enabled product.” Compensaciones / Beneficios • FlexBase program, • remote/hybrid/work-from-home options, • health and well-being benefits, • financial planning benefits, • life beyond work support Responsabilidades • Operate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal tooling, • Enhance platform reliability, observability, security, performance, and cost efficiency, • Provide guidance to engineers to increase confidence in service performance, • Lead complex production investigations and drive long-term improvements, • Leverage LLMs and AI-driven automation to auto-remediate incidents and streamline operations, • Partner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problems, • Participate in on-call rotations, guiding restoration and repair of service-impacting issues Requisitos principales • 3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systems, • Design and implement monitoring/observability strategy using Prometheus and Grafana, • Production experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on Linux, • Exceptional troubleshooting across network, system, applications, and databases, • Experience with GitOps, CI/CD, and Infrastructure as Code, • Scripting/programming in Python, Go, and Bash, • Experience using AI tools in operations and proposing automation initiatives, • Technical leadership and ownership in cross-team initiatives, • technical leadership, • ownership, • cross-functional collaboration, • Kubernetes, • Docker, • Helm