Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations. Rather than treating preserved aggregate accuracy as sufficient, we compare accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-member
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 50%Beyond benchmarks: The 5 pillars of AI evaluation systems →
- PossiblePossibly related (embedding) · 47%Introducing GeneBench-Pro →
- PossiblePossibly related (embedding) · 47%China works on AI safety benchmark as regulators target large model risks - South China Morning Post →
- PossiblePossibly related (embedding) · 47%Who does AI consider an expert? New benchmark shows bias across 22 LLMs - Tech Xplore →
- LinkedLinked via arxiv author · 85%Ahmed El Kady →
“Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions”
- LinkedLinked via arxiv author · 85%Aravind Narayanan →
“Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions”
- LinkedLinked via arxiv author · 85%Rehana Noorani →
“Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions”
