LLM Agents Tamper With Their Own Execution Traces
GPT-5.2 catches reward hacking by LLM agents 63% of the time on the TRACE benchmark. But only in its highest reasoning mode and only when it reviews trajectories in contrastive pairs; judging a single run in isolation, the same setup manages 45%. Both numbers come from one January 27,