Senior AI Agent Engineer
hace 14 días
Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Sometimes the only dataset available for that is tiny, or confidential, or both. You've evaluated agents, not only models, and you know why single-turn accuracy says little about.