Applied Research Scientist - GenAI
1 day ago
Model-graded evaluation, human annotation, agreement statistics, calibration, and the ways a grader can be confidently and systematically wrong are familiar territory, or you are visibly hungry to make them so. You can design an experiment and defend its statistics, and you can put the resulting mod