← back
arXivHan Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu, Yue ZhangFri, Oct 2, 2026, 4:57 AM PDT
score 14.9

Why AI Distillation Can Produce Repetitive, Overlong Output

Original: Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

Source: arxiv.org ↗

Writing ELI5 summary…