Site Reliability Engineer
hace 5 días
New York
What is Aisle?\n Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened. \n Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact. \n This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time. \n Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment. \n We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis. \n You'll win here if…\n\n • You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio\n, • Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system\n, • You’re comfortable moving quickly, shipping improvements, and iterating in production\n, • AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster\n, • You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)\n, • You have experience building or orchestrating AI/agent workflows - in production or through serious side projects\n\n About your role\n Reliability & Infrastructure \n\n • Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel\n, • Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards\n, • Manage IAM, service accounts, and security best practices across our cloud environment\n, • Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements\n\n Core Infrastructure & State \n\n • Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows\n, • Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments\n, • Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution\n, • Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)\n, • Automate infrastructure provisioning, deployments, and operational workflows\n\n AI & Next-Gen Tooling: Agent Ops \n\n • Build agent operations infrastructure that enables AI agents to run safely and reliably in production\n, • Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry\n, • Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop\n, • Own visibility into AI usage, reliability, and spend as our agent footprint scales\n\n Cross-functional Impact \n\n • Partner closely with engineering and product teams to maintain reliability without slowing development velocity\n, • Act as a force multiplier across the team — helping engineers ship faster and more safely \nAbout your skillsMust haves\n\n • 4+ years in SRE, DevOps, or infrastructure/platform engineering\n, • Strong, hands-on experience with a major cloud platform (preferable GCP)\n, • Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)\n, • Solid understanding of IAM, security, and cloud best practices\n, • Experience with observability tools like Datadog\n, • Familiarity with Node.js environments\n, • AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.\n, • Experience building or orchestrating AI/agent workflows (work or serious personal projects)\n, • High ownership, strong curiosity, and a bias toward action\nNice to haves\n\n, • Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls\n, • Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus\n, • Hands-on experience with GCP Workflows for orchestration\n, • Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling\n, • Experience with Vercel deployment and edge computing\n, • Familiarity with the k8s ecosystem\n, • Familiarity with Redis and BullMQ\n, • Understanding of SOC 2 compliance requirements and implementation\n, • Previous experience in a high-growth startup environment\n, • Previous backend engineering experience to bridge the gap between infrastructure and code\nAbout the stack\n\n, • Cloud: Google Cloud Platform (GCP)\n, • Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes\n, • Database: PostgreSQL with pgbouncer, Prisma\n, • Observability: Datadog\n, • Runtime: Node.js, TypeScript\n, • Deployment: Vercel, GCP\n, • Frontend: React, Next.js, TypeScript\n, • Backend: TypeScript, PostgreSQL, Next.js\n