Megadose AI progress, ranked and analyzed.

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

· HF Daily Papers ·
CalibForge builds terminal-agent training tasks by calibrating them against actual solver behavior.

The system revises candidate tasks until they sit in a solver-relative “learnable zone,” either by provoking disagreement across solvers or by making a strong solver pass while a weaker one fails. The authors report a collection of 5,431 calibrated terminal tasks. In ablations, those calibration strategies outperform authoring plus validation alone and ordinary single-solver feedback. Models trained on the full set reached 32.58% and 47.57% on Terminal-Bench 2.0, with the largest gains also reported on SWE-bench Pro and Doc2Repo. HF Daily Papers' note

score 5

Categories: Research