Megadose AI progress, ranked and analyzed.

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

· ArXiv · AI/CL/LG ·
The paper argues that only some calls in a multi-turn tool workflow are actually worth training.

Critical-State RL tests candidate model calls to see whether their local reward reflects the action’s effect on task success, rather than later randomness. It then uses nested sampling to isolate action-dependent variation and trains only the selected states with a contextual-bandit setup. On BFCL v4, those diagnostics picked different turns for missing-function and missing-argument tasks; training those turns improved results, including about 14 points on the missing-function task. ArXiv · AI/CL/LG's note

score 5

Categories: Research