← back
arXivOrion Reblitz-RichardsonTue, Oct 6, 2026, 9:52 AM PDT
score 16.7

AI Models Sometimes Act Against Their Own Stated Moral Judgment

Original: Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment

Source: arxiv.org ↗

Writing ELI5 summary…