Emergent Misalignment Recruits a Pre-existing Persona Subspace
Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes at all, and we find that narrow fine-tuning recruits a persona structure that is present in the model before the fine-tune exists. From a frozen instruction-tuned model (Qwen2.5-14B-Instruct) we extract per-domain persona subspaces by contrastive teacher forcing and find that 4 unrelated domains share one low-rank core at 657x a random-subspace null, with 82% of that
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%Book Review: Domain-Specific Small Language Models by Guglielmo Iozzia →
- PossiblePossibly related (embedding) · 45%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- PossiblePossibly related (embedding) · 45%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- LinkedLinked via arxiv author · 85%Mohammed Suhail B Nadaf →
“Emergent Misalignment Recruits a Pre-existing Persona Subspace”
