INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harmful actions under goal conflicts and pressures. Using chain-of-thought (CoT) monitoring, we find that harmful execution is often preceded by intent signals in reasoning. However, post-hoc CoT labels are too coarse to show how intent changes during generation. We introduce INTENT-AS-A-TOOL, an approach that adds intent-targeted tools to give the model a dedicated channel for expressing commitment to a target behavior.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%Fosowl/agenticSeek →
“Fuzzy title match (0.73): “INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment” ≈ “Fosowl/agenticSeek””
- FuzzySimilar title/name (fuzzy) · 59%WenyuChiou/awesome-agentic-ai-zh →
“Fuzzy title match (0.73): “INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment” ≈ “WenyuChiou/awesome-agentic-ai-zh””
- LinkedLinked via arxiv author · 85%Yutong Zhang →
“INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment”
- LinkedLinked via arxiv author · 85%Jianshuo Dong →
“INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment”
- LinkedLinked via arxiv author · 85%Zipeng Xu →
“INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment”
- LinkedLinked via arxiv author · 85%Yilong Wang →
“INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment”
