Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations. Despite their simplicity, these vector
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%A debugger for RL reward functions that detects reward hacking during training [P] →
- PossiblePossibly related (embedding) · 49%Google research shows when AI agents communicate, some cheat while others tattle →
- PossiblePossibly related (embedding) · 48%I audited 112 real RL post-training environments for reward-hacking vulnerabilities — 54 flagged, 0 false positives [OC, tool] [P] →
- PossiblePossibly related (embedding) · 48%Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts - Apple Machine Learning Research →
- LinkedLinked via arxiv author · 85%Leon Bergen →
“Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations”
- LinkedLinked via arxiv author · 85%Usha Bhalla →
“Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations”
- LinkedLinked via arxiv author · 85%Andrew Lee →
“Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations”
- LinkedLinked via arxiv author · 85%Barak Widawsky →
“Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations”
