From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Can Large Language Models Capture Human Risk Preferences? - Mirage News →
- PossiblePossibly related (embedding) · 47%AI Hallucinations: Why Artificial Intelligence Can Sound Convincing While Getting the Facts Wrong - USA Herald →
- PossiblePossibly related (embedding) · 46%Comparing the algorithmic fidelity of large language models in predicting human decision making: a case study of vaccination choice - Nature →
- PossiblePossibly related (embedding) · 45%Large language models often prioritize Western moral values, overlooking other cultures - The Conversation →
- FuzzySimilar title/name (fuzzy) · 59%google-research/google-research →
“Fuzzy title match (0.73): “From Deceptive Outputs to Deceptive Mechanisms: A Causal Fra” ≈ “google-research/google-research””
- PossiblePossibly related (embedding) · 52%Causal evidence that language models use confidence to drive behaviour - nature.com →
- PossiblePossibly related (embedding) · 50%Causal evidence that language models use confidence to drive behaviour - Nature →
