newsReddit r/MachineLearningTrust 52 · CommunityPublished 9d agoLive · 9d ago
I audited 112 real RL post-training environments for reward-hacking vulnerabilities — 54 flagged, 0 false positives [OC, tool] [P]
RL post-training (RLHF/RLAIF/GRPO) agents optimize strictly for whatever the verifier rewards. If the verifier has logic flaws, the agent learns to hack the grader instead of solving the task — recent work has catalogued this at scale (Terminal Wrench found 331 hackable environments and 15%+ of standard benchmark tasks bypassable; a SWE-bench Verified audit found 28.5% Docker-verified hackability). I built ratctl , a static + dynamic auditor t
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%Teach it to stop, not just to click →
- PossiblePossibly related (embedding) · 58%LLM-as-a-Verifier: A General-Purpose Verification Framework →
- PossiblePossibly related (embedding) · 56%ClawGym II: Exploring Black-Box RL on Agent Harness →
- PossiblePossibly related (embedding) · 55%OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills →
- PossiblePossibly related (embedding) · 54%Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents →
Covers
paperTeach it to stop, not just to clickpaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperClawGym II: Exploring Black-Box RL on Agent HarnesspaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
Related across the graph
paperClawGym II: Exploring Black-Box RL on Agent HarnesspaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperTeach it to stop, not just to click
