Applied Research Scientist - GenAI
19 hours ago
Model-graded evaluation, human annotation, agreement statistics, calibration, and the ways a grader can be confidently and systematically wrong are familiar territory, or you are visibly hungry to make them so. You can design an experiment and defend its statistics, and you can put the resulting mod