Senior AI Agent Engineer
hace 14 días
Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt. You've used at least one LLM evaluation framework, in-house tooling included. You've built internal tooling that non-engineers used on their own to label and review model output.