arXivZejun Xie, Xintong Li, Guang Wang, Desheng ZhangMon, Aug 3, 2026, 9:30 AM PDT
score 17.0
Combining human rankings with AI scores improves evaluation when true answers are unknown
Original: Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
Source: arxiv.org ↗
Writing ELI5 summary…