← back
arXivZejun Xie, Xintong Li, Guang Wang, Desheng ZhangMon, Aug 3, 2026, 9:30 AM PDT
score 17.0

Combining human rankings with AI scores improves evaluation when true answers are unknown

Original: Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

Source: arxiv.org

Writing ELI5 summary…