How Language Models Organize and Structure Moral Knowledge
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%Large language models often prioritize Western moral values, overlooking other cultures - The Conversation →
- PossiblePossibly related (embedding) · 47%Large language models could democratize data use—but only if the foundations are in place - World Bank Blogs →
- LinkedLinked via arxiv author · 85%Orion Reblitz-Richardson →
“How Language Models Organize and Structure Moral Knowledge”
