Senior AI Agent Engineer
hace 14 días
Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt. You've evaluated agents, not only models, and you know why single-turn accuracy says little about. You've built internal tooling that non-engineers used on their own to label and review model output.