On the Threat Model of Weird Generalization and Emergent Misalignment
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive th
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- PossiblePossibly related (embedding) · 47%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- PossiblePossibly related (embedding) · 47%The shrinking landscape of linguistic diversity in the age of large language models - Nature →
- PossiblePossibly related (embedding) · 46%Alignment →
- PossiblePossibly related (embedding) · 45%Does anyone have a name for that subtle "Sameness" creeping into model outputs lately? [R] →
- LinkedLinked via arxiv author · 85%Mark Dredze →
“On the Threat Model of Weird Generalization and Emergent Misalignment”
- LinkedLinked via arxiv author · 85%Miriam Wanner →
“On the Threat Model of Weird Generalization and Emergent Misalignment”
- LinkedLinked via arxiv author · 85%William Walden →
“On the Threat Model of Weird Generalization and Emergent Misalignment”
