Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safety of LLM reasoning. Cross-architecture, cross-scale, and cross-dataset checks show that RIM does not always emerge. Previous work attributed RIM to neuron-level entanglement, but did not identify the geometry of the representation space underlying this entanglement or propose a training-time fix. We provide both: a representation-space analysis
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%I tried to give an LLM room to think. This is where it led. [P] →
- PossiblePossibly related (embedding) · 47%A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models - Apple Machine Learning Research →
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Yipeng Zhao →
“Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty”
- LinkedLinked via arxiv author · 85%Qishun Yang →
“Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty”
- LinkedLinked via arxiv author · 85%Shenzhe Zhu →
“Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty”
- LinkedLinked via arxiv author · 85%Shu Yang →
“Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty”
