Senior AI Agent Engineer
hace 13 días
Stand up eval suites using various evaluation frameworks and tooling, included but not limited to. Sometimes the only dataset available for that is tiny, or confidential, or both. Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt.