Sr Forward Deployment Engineer
hace 2 días
Chicago
Senior Forward Deployed GenAI Engineer \n Hybrid Opportunity- Chicago (3 days a week at clients office) \n Re-location provided. \n \n About Avanta \n Avanta helps ambitious small and mid-market companies build AI intelligence that changes how work gets done—in weeks, not quarters. We design and deliver practical AI systems from model selection through production deployment, and we build them so the customer can understand, operate, and own the outcome. \n Our work spans agentic solutions, AI enablement, AI-native software delivery, business transformation, small language models, strategy, governance, adoption, measurement, and executive advisory. \n \n Role Summary \n You will independently design and deliver production generative AI applications while balancing customer value, technical quality, security, cost, adoption, and operational ownership. You will mentor associate engineers and serve as a trusted technical partner to customer teams. \n \n What You Will Do \n\n • Lead the design and implementation of production-grade generative AI applications and customer workstreams.\n, • Build RAG systems across structured, semi-structured, and unstructured data, including authorization-aware retrieval and citations.\n, • Design agentic workflows coordinating multiple tools, services, business rules, and human approval boundaries.\n, • Evaluate proprietary, open-weight, multimodal, embedding, reranking, and specialized models for specific business needs.\n, • Benchmark quality, accuracy, latency, throughput, cost, safety, and operational complexity before recommending a model or architecture.\n, • Create golden datasets, scoring rubrics, regression tests, and repeatable evaluation pipelines.\n, • Diagnose failures across ingestion, parsing, chunking, embeddings, retrieval, reranking, prompts, tool execution, orchestration, and output generation.\n, • Implement human review, approval, escalation, and rollback mechanisms for higher-impact actions.\n, • Design for multitenancy, resiliency, availability, graceful degradation, and defined service-level objectives.\n, • Integrate AI applications with enterprise systems, data platforms, workflow tools, identity services, and security controls.\n, • Build production monitoring for model quality, retrieval effectiveness, tool execution, latency, token usage, cost, errors, and user behavior.\n, • Improve systems through caching, batching, context reduction, prompt optimization, model routing, provider abstraction, and architectural changes.\n, • Build deployment pipelines, infrastructure automation, automated testing, and release-management practices.\n, • Support production incidents, lead root-cause analysis for owned workstreams, and implement durable corrective actions.\n, • Convert customer implementations into templates, connectors, evaluation packs, runbooks, accelerators, and reference architectures.\n, • Mentor associate engineers and provide code, design, and customer-communication feedback.\n\n \n Customer-Facing Responsibilities \n\n • You will lead technical discovery with customer engineers, business owners, security teams, operators, and data stakeholders.\n, • You'll translate business problems into measurable AI use cases and technical requirements, facilitate design workshops, create implementation plans, coordinate across customer teams, communicate risks, and help customer engineers understand and operate the delivered solution.\n, • You may also support pre-sales technical discovery when needed, while remaining accountable for hands-on delivery.\n\n \n Technical Responsibilities \n\n • Design cloud-native services using API, distributed-system, asynchronous, and event-driven patterns.\n, • Apply hybrid retrieval, metadata filtering, query rewriting, reranking, retrieval evaluation, and graph-based retrieval where appropriate.\n, • Design agent state, memory, tool contracts, authorization, failure recovery, and deterministic controls.\n, • Establish prompt/configuration versioning, automated regression testing, experiment tracking, and release gates.\n, • Implement identity-aware retrieval and least-privilege access.\n, • Protect systems from prompt injection, data leakage, unsafe tool use, malicious input, and invalid structured output.\n, • Apply containers, Kubernetes or serverless patterns, infrastructure as code, and automated CI/CD.\n, • Trace and observe the complete application path across models, retrieval, agents, tools, data, and infrastructure.\n, • Optimize model and application cost, latency, throughput, and reliability.\n, • Define runbooks, incident procedures, support boundaries, SLOs, and ownership before go-live.\n\n \n Required Qualifications \n\n • Approximately 4–7 years of relevant engineering experience, including meaningful ownership of production software, enterprise AI, ML, search, data, or automation systems.\n, • Strong production software-engineering ability in Python, plus competence in at least one additional enterprise language or framework.\n, • Experience designing and operating APIs, cloud services, data pipelines, integrations, automated tests, observability, and release pipelines.\n, • Hands-on experience taking a GenAI or ML solution beyond prototype into production.\n, • Practical expertise in RAG, model evaluation, agent tools, security controls, cost management, and production troubleshooting.\n, • Ability to lead technical discovery and explain trade-offs to business leaders, customer engineers, and security stakeholders.\n, • Evidence of independent ownership, sound judgment, mentorship, and disciplined delivery in ambiguous environments.\n\n \n Representative Technology Stack \n Languages: Advanced Python, TypeScript/JavaScript, Java, C#, Go, SQL, FastAPI, Node.js, React. \n AI/Models: Amazon Bedrock, SageMaker, Azure OpenAI, AI Foundry, Google Vertex AI, OpenAI, Anthropic, Hugging Face, open-weight models. \n Agents: LangGraph, LangChain, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, Haystack. \n Retrieval/Data: PostgreSQL/pgvector, OpenSearch/Elasticsearch, Pinecone, Weaviate, Milvus, Qdrant, Redis, Neo4j, warehouses, lakehouses and streaming platforms. \n Evaluation/Observability: MLflow, Weights & Biases, LangSmith, Arize Phoenix, OpenTelemetry, Prometheus, Grafana. \n Infrastructure: Docker, Kubernetes, Terraform, CI/CD, Argo CD, serverless, API gateways and model gateways. \n Security: IAM, RBAC/ABAC, encryption, DLP, audit logging, policy engines, guardrails, lineage and governance tooling. \n \n \n \n What Success Looks Like \n You'll independently own customer workstreams from discovery through production, improve AI system quality, reliability, latency and cost using evidence, establish repeatable evaluation and observability practices, turn prototypes into maintainable production systems, mentor engineers, and convert successful implementations into reusable Avanta capabilities. \n \n Important note: This is not a presentation-only Sales Engineer role. It remains a hands-on delivery role. \n