Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
The paper reports an 8.95-point macro-average gain from learned harness-building rules over no-skill construction.
The authors frame the setup as test-time AI-for-AI: a Builder model improves the execution environment for a Target model without changing either model’s weights. Its “Meta-Skill” bank captures when the Target needs support and what resources the harness should provide. Learned from development-set execution feedback, that bank is then frozen and used on unseen tasks. The reported gains hold across Harness-Bench and NewtonBench, including a 12.02-point advantage over giving the same bank directly to the Target. HF Daily Papers' note
The authors frame the setup as test-time AI-for-AI: a Builder model improves the execution environment for a Target model without changing either model’s weights. Its “Meta-Skill” bank captures when the Target needs support and what resources the harness should provide. Learned from development-set execution feedback, that bank is then frozen and used on unseen tasks. The reported gains hold across Harness-Bench and NewtonBench, including a 12.02-point advantage over giving the same bank directly to the Target. HF Daily Papers' note
score 5