Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection". We initially discover this attack in the multiple-choice question-answering monitorability setting proposed by Lanham et al. (2023), using the investigator-agent elicitation framework of
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Copilot tricked into telling reseachers how to hack itself →
- PossiblePossibly related (embedding) · 54%Chain-of-Thought Spoofing Targets Reasoning AI Models - Hackaday →
- PossiblePossibly related (embedding) · 54%Google research shows when AI agents communicate, some cheat while others tattle →
- PossiblePossibly related (embedding) · 52%Adversarial Code Review: ICML Study Shows 3 AI Agents With Structured Disagreement Beat 5-Agent Teams - Intelligent Living →
- LinkedLinked via arxiv author · 85%Keertana Chidambaram →
“Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection”
- LinkedLinked via arxiv author · 85%Andrew Ilyas →
“Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection”
- LinkedLinked via arxiv author · 85%Vasilis Syrgkanis →
“Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection”
