Senior AI Agent Engineer
hace 9 días
You've built internal tooling that non-engineers used on their own to label and review model output. Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt.