Senior AI Agent Engineer
13 days ago
Stand up eval suites using various evaluation frameworks and tooling, included but not limited to. Build the citations and confidence. Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt.