Senior AI Agent Engineer
6 days ago
Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Sometimes the only dataset available for that is tiny, or confidential, or both. You've evaluated agents, not only models, and you know why single-turn accuracy says little about.