AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
A stronger model built test-time harnesses that lifted weaker models from 0.49 to 0.91 average performance without retraining.
The paper tests “strong-to-weak scaffolding” on four Theory-of-Mind benchmarks. Builder models used 5% of the data to refine harnesses, then ran the finished harness on the full test set. The gains came mostly from deterministic code, benchmark-specific routing, and strict answer formatting, not from making the weaker model reason longer. Weaker target models benefited most. ArXiv · AI/CL/LG's note
The paper tests “strong-to-weak scaffolding” on four Theory-of-Mind benchmarks. Builder models used 5% of the data to refine harnesses, then ran the finished harness on the full test set. The gains came mostly from deterministic code, benchmark-specific routing, and strict answer formatting, not from making the weaker model reason longer. Weaker target models benefited most. ArXiv · AI/CL/LG's note
score 6