The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits
Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the publish
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Cutting RAG inference costs 6x starts with deciding what never reaches the LLM →
- PossiblePossibly related (embedding) · 47%AI is more likely than humans to form biases when hiring →
- PossiblePossibly related (embedding) · 46%A Psychometric Comparison of Faculty-Authored and Large Language Model-Generated Multiple-Choice Questions in Endodontics - Cureus →
- LinkedLinked via arxiv author · 85%Siddharth Vohra →
“The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits”
- LinkedLinked via arxiv author · 85%Manikandan Ravikiran →
“The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits”
