Senior AI Agent Engineer
hace 14 días
Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt. You've used at least one LLM evaluation framework, in-house tooling included. You can tell a real regression from noise, and design an experiment that answers the question.