Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
Task-CoEvolve claims full-set search performance with 80% fewer optimization evaluations.
The paper targets LLM agent harness optimization, where code around the model is rewritten based on validation results while model weights stay fixed. Its method samples validation tasks where candidate harnesses disagree, treating those as more useful than tasks everyone solves or fails. It estimates full-set scores from the sampled subset by accounting for sampling probabilities. Experiments on online text classification and Terminal-Bench 2.1 beat fixed-subset baselines and matched full-set search’s final performance. ArXiv · AI/CL/LG's note
The paper targets LLM agent harness optimization, where code around the model is rewritten based on validation results while model weights stay fixed. Its method samples validation tasks where candidate harnesses disagree, treating those as more useful than tasks everyone solves or fails. It estimates full-set scores from the sampled subset by accounting for sampling probabilities. Experiments on online text classification and Terminal-Bench 2.1 beat fixed-subset baselines and matched full-set search’s final performance. ArXiv · AI/CL/LG's note
score 5