Applied Research Scientist - GenAI
1 day ago
Model-graded evaluation, human annotation, agreement statistics, calibration, and the ways a grader can be confidently and systematically wrong are familiar territory, or you are visibly hungry to make them so. The limit on every self-improving system is the quality of the signal that grades it, and