Phantom Gains: Auditing Self-Improvement Against a Measured Null
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which inverts a reported finding when its control is absent. Several are standard practice. A ledger built on a single greedy decode manufactures capability change
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Competence Gate: gating tool-use on a small model's internal confidence signal instead of its verbalised one — Qwen3.5-4B, open weights [P] →
- LinkedLinked via arxiv author · 85%Yicheng Xu →
“Phantom Gains: Auditing Self-Improvement Against a Measured Null”
- LinkedLinked via arxiv author · 85%Nan Yang →
“Phantom Gains: Auditing Self-Improvement Against a Measured Null”
- LinkedLinked via arxiv author · 85%Liming Chen →
“Phantom Gains: Auditing Self-Improvement Against a Measured Null”
- LinkedLinked via arxiv author · 85%M-Tahar Kechadi →
“Phantom Gains: Auditing Self-Improvement Against a Measured Null”
