Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign) →
- PossiblePossibly related (embedding) · 48%Implicit-bias-like patterns in reasoning models →
- FuzzyOverlapping authors or contributors · 62%open-webui/open-webui →
“Shared author/contributor keys: nguyen”
- LinkedLinked via arxiv author · 85%Hoang Cuong Nguyen →
“Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness”
- LinkedLinked via arxiv author · 85%Mark Dras →
“Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness”
- LinkedLinked via arxiv author · 85%Usman Naseem →
“Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness”
