Phantom Gains: Auditing Self-Improvement Against a Measured Null
The paper says common transition-level audits can report self-improvement even when the model was frozen.
The authors ran Qwen3-8B LoRA self-training against a frozen control put through the same pipeline and found seven measurement failures that could flip conclusions without that control. A single greedy-decode ledger produced apparent capability changes on the untrained model, largely from inference batching, and a threshold fix still left a non-zero null. Their replacement uses per-problem exact tests against pooled baseline replicates with false-discovery-rate control, detecting no held-out self-training gains. External distillation improved problems the base model rarely solved, while the tested self-training variants did not and also corrupted some baseline-solved problems above the measured floor. ArXiv · AI/CL/LG's note
The authors ran Qwen3-8B LoRA self-training against a frozen control put through the same pipeline and found seven measurement failures that could flip conclusions without that control. A single greedy-decode ledger produced apparent capability changes on the untrained model, largely from inference batching, and a threshold fix still left a non-zero null. Their replacement uses per-problem exact tests against pooled baseline replicates with false-discovery-rate control, detecting no held-out self-training gains. External distillation improved problems the base model rarely solved, while the tested self-training variants did not and also corrupted some baseline-solved problems above the measured floor. ArXiv · AI/CL/LG's note
score 4