How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI
Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- PossiblePossibly related (embedding) · 49%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5 →
“Shared author/contributor keys: choi”
- FuzzySimilar title/name (fuzzy) · 59%microsoft/semantic-kernel →
“Fuzzy title match (0.73): “How Much Human Label Variation Does Formal Semantic Structur” ≈ “microsoft/semantic-kernel””
- FuzzySimilar title/name (fuzzy) · 59%vllm-project/semantic-router →
“Fuzzy title match (0.73): “How Much Human Label Variation Does Formal Semantic Structur” ≈ “vllm-project/semantic-router””
- LinkedLinked via arxiv author · 85%Haram Choi →
“How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in N”
