Validity of LLMs as data annotators: AMALIA on authority
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory o
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%The Large Language Model (LLM) understands the world through human texts. A world model is needed f.. - 매일경제 →
- PossiblePossibly related (embedding) · 47%sileod/llm-theory-of-mind →
- PossiblePossibly related (embedding) · 47%VectorPeak/LLM-Wiki →
- PossiblePossibly related (embedding) · 45%Jeryi-Sun/LLM-and-Law →
- LinkedLinked via arxiv author · 85%Manuel Pita →
“Validity of LLMs as data annotators: AMALIA on authority”
- PossiblePossibly related (embedding) · 48%Large language models often prioritize Western moral values, overlooking other cultures - The Conversation →
