Harness Learning Enables Generalizable Test-Time Adaptation
The paper trains a model to revise an agent’s execution harness at test time, using run feedback instead of changing model weights.
The authors define the harness as the program around a language model: calls, tools, and information flow. Their “harness learning” setup uses reinforcement learning to train a proposer that edits that harness after seeing execution feedback. In reasoning and multi-hop QA experiments, revised harnesses improved, and the adaptation carried over to unseen tasks. The paper says single-revision policies could keep improving over multiple rounds, while sequence training helped unevenly across settings. ArXiv · AI/CL/LG's note
The authors define the harness as the program around a language model: calls, tools, and information flow. Their “harness learning” setup uses reinforcement learning to train a proposer that edits that harness after seeing execution feedback. In reasoning and multi-hop QA experiments, revised harnesses improved, and the adaptation carried over to unseen tasks. The paper says single-revision policies could keep improving over multiple rounds, while sequence training helped unevenly across settings. ArXiv · AI/CL/LG's note
score 6