Megadose AI progress, ranked and analyzed.

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

· ArXiv · AI/CL/LG ·
MATCH trains tool use around the model’s moving capability boundary, then withholds lower-level credit when the tool choice is wrong.

The paper introduces MACL, a curriculum that refreshes sample difficulty from reward signals and selects cases near the policy’s current edge, plus harder top-k examples. Its HTGR reward scores tool name, argument key, and argument value in order, granting each layer only if the prior one is correct. The same gated reward drives both GRPO optimization and the curriculum update. On API-Bank and BFCL V3, the authors report 72.19% and 62.87% overall accuracy, above their supervised and RL baselines. ArXiv · AI/CL/LG's note

score 5

Categories: Research