Megadose AI progress, ranked and analyzed.

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

· ArXiv · AI/CL/LG ·
A stronger model built test-time harnesses that lifted weaker models from 0.49 to 0.91 average performance without retraining.

The paper tests “strong-to-weak scaffolding” on four Theory-of-Mind benchmarks. Builder models used 5% of the data to refine harnesses, then ran the finished harness on the full test set. The gains came mostly from deterministic code, benchmark-specific routing, and strict answer formatting, not from making the weaker model reason longer. Weaker target models benefited most. ArXiv · AI/CL/LG's note

score 6

Categories: Research