Senior GenAI Platform Engineer
10 days ago
Madrid
ph3About Pleo /h3 pMessy spend management is tricky business. And tedious processes are a lose‑lose situation for all involved, not just finance. At Pleo, we're changing that. We build spend solutions that make managing money seamless, empowering, and surprisingly effective for finance teams and employees alike - with a vision to help all businesses ‘go beyond’. /p pThe word “Pleo” actually means “more than you’d expect”, and living by that mantra has been the secret to our success over the last 10 years. Now, we’re at a pivotal moment in our journey; every move we make has a direct impact on our 40,000+ customers, our business, and our collective success. We need people who take pride in uncovering customer needs, who turn complex problems into simple solutions, challenge the way things are done (respectfully), and always aim high. With great ambitions driving us forward, we can’t say we’ve got this whole thing figured out. And frankly, that’s half the fun! What we can say is that we’re a driven, progressive, and, importantly, a kind bunch of 850+ people from over 100 nationalities, all committed to delivering the future of business spending, together. /p h3The Role /h3 pPleo is investing heavily in AI‑powered features across the product. You will be part of the GenAI Core team which is responsible for the horizontal platform infrastructure that makes this possible. They look after LLM routing, MCP servers, vector search infrastructure, evaluation frameworks, and agentic tooling. /p pThis is a hands‑on backend/platform engineering role at the intersection of distributed systems and modern AI engineering. You will help design, build, and operate the shared AI infrastructure used by product teams across Pleo, with a strong focus on reliability, observability, security, and developer experience. /p h3Reporting and Collaboration /h3 pYou’ll be reporting to the Engineering Manager for the GenAI Platform team and will be working closely with senior and staff Engineers in GenAI Core. You will also collaborate with Applied AI Engineers, Data Scientists, and product engineering teams across the business. /p h3What You’ll Be Doing /h3 ul liDesign, build, and operate core GenAI platform components used by product teams at Pleo, including LLM routing gateway, vector search and RAG infrastructure, tool registry and MCP gateway, AI observability and evaluation tooling (tracing LLM calls, supporting human and automated evaluation, detecting drift, and tracking costs) and infrastructure for multi‑step, long‑running agentic workflows. /li liOwn production‑quality delivery of platform features, from design through rollout, monitoring, and follow‑up. /li liContribute to resilient system design: sensible APIs, failure handling, rate limiting, retries, idempotency, and safe change management. /li liImprove reliability and observability through metrics, dashboards, alerting, incident follow‑ups, and operational improvements. /li liPartner with Applied AI Engineers and product teams to understand platform needs and help them build AI‑powered features safely. /li liBuild internal SDKs, templates, and guardrails that let product engineers build AI features without needing deep infrastructure expertise. /li liSupport other engineers through pairing, code reviews, technical feedback, and clear documentation. /li liHelp evaluate build‑vs‑buy decisions in the rapidly evolving LLMOps tooling landscape. /li /ul h3What You Bring /h3 ul liStrong backend/systems engineering background, with experience building and operating production services with reliability and observability requirements. /li liExperience designing and delivering shared platform or infrastructure components used by multiple teams. /li liStrong production ownership: monitoring, alerting, incident response, debugging, and post‑incident learning. /li liDistributed systems fundamentals, including async workflows, idempotency, consistency tradeoffs, and designing for failure. /li liHands‑on experience with LLM APIs or strong interest in learning their production failure modes: rate limits, context windows, multi‑vendor routing, latency variance, and cost control. /li liSecurity mindset for AI systems, including prompt injection risks, PII in logs, data leakage, and safe credential handling. /li liStrong programming experience in either a JVM‑based language or Python. We operate a polyglot platform with components written in both Kotlin and Python, and you’ll be expected to contribute to both. /li liClear communication and collaboration skills, especially when working with product teams and other engineers to turn ambiguous platform needs into practical solutions. /li /ul h3Fit Assessment /h3 pbGood Fit: /b /p ul liYou are passionate about developer experience and get excited about being a force multiplier for engineering teams. /li liYou have moved past prototyping and have a deep understanding of the realities of LLMOps, data retrieval, prompt and context engineering, as well as model evaluation in production. /li liYou understand both the engineering and the data side of things, and are comfortable switching languages or technologies to achieve your goals. /li /ul pbNot a good fit: /b /p ul liYou are primarily interested in model research or algorithm development. This role is about building the tooling that enables product teams to ship AI features to production. /li liYou prefer building customer‑facing features. Your key users will be other Pleo Engineers. /li liYou cannot explain AI trade‑offs clearly to non‑technical stakeholders. You will regularly work with Product Managers, Designers, and business leaders who need to understand what is and isn’t possible. /li /ul h3Development Path /h3 ul liDevelop a clear picture of how AI features are currently being built at Pleo and where the biggest infrastructure bottlenecks are. /li liTake ownership of a core platform component such as the LLM gateway, RAG infrastructure, MCP gateway, or evaluation framework, and improve its reliability, observability, or developer experience. /li liDeliver production‑ready improvements with clear rollout plans, monitoring, and operational documentation. /li liPartner with Applied AI Engineers and product teams to identify platform investments that would accelerate their work. /li liContribute to Pleo's internal standards for AI feature development: how we evaluate quality, manage prompts, and monitor production AI systems. /li /ul h3Location /h3 pWe can hire on a remote, hybrid or in‑person set‑up in any of the locations listed on the advert but you will need to be physically based in the country of your choice with a valid right to work. We are unable to offer visa sponsorship for this role. /p h3Benefits /h3 ul liYour own Pleo card (no more out‑of‑pocket spending!) /li liLunch is on us for your work days – enjoy catered meals or receive a lunch allowance based on your local office /li liComprehensive private healthcare – depending on your location, coverage options include Vitality, Alan or Médis /li liWe offer 25 days of holiday + your public holidays /li liHybrid and fully remote working options for our team /li liAccess to free mental health and well‑being support via MyndUp /li liPaid parental leave – we want to support families and help you feel you don’t have to compromise your family due to work /li /ul /p #J-18808-Ljbffr