Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
Combining human judges and AI scores for better decisions without ground truth
When assessing subjective things like essay quality or medical images, humans and AI models each have blind spots: humans disagree with each other, while AI learns from incomplete information. Researchers developed a two-stage method that first finds consensus among human judges, then uses that consensus to correct and calibrate AI scores. The approach outperformed using either humans or AI alone across multiple real-world tests.
Many high-stakes decisions—hiring, healthcare, academic grading—rely on assessment where there's no perfect right answer to check against. This method lets organizations systematically improve their judgment calls by combining what humans do well (comparing options) with what AI does well (spotting patterns). The theoretical guarantees mean practitioners can trust the approach even when human agreement is messy or incomplete.