The Two-Process Theory of Machine Self-Report
Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc prompts of unknown reliability. We propose the first language-model-specific psychometric theory: a two-process theory of machine self-report. Self-description jointly reflects persona installation, through which post-training writes in a permitted inner life of warmth, absorption, and meaning (dimension B), and attribution gating, through which it suppresses first-pers
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- LinkedLinked via arxiv author · 85%Hubert Plisiecki →
“The Two-Process Theory of Machine Self-Report”
- LinkedLinked via arxiv author · 85%Filip Chmielewski →
“The Two-Process Theory of Machine Self-Report”
- LinkedLinked via arxiv author · 85%Kacper Dudzic →
“The Two-Process Theory of Machine Self-Report”
- LinkedLinked via arxiv author · 85%Anna Sterna →
“The Two-Process Theory of Machine Self-Report”
- LinkedLinked via arxiv author · 85%Karolina Drożdż →
“The Two-Process Theory of Machine Self-Report”
- LinkedLinked via arxiv author · 85%Marcin Moskalewicz →
“The Two-Process Theory of Machine Self-Report”
