Senior AI Agent Engineer
6 days ago
Stand up eval suites using various evaluation frameworks and tooling, included but not limited to. Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt.