Looped Language Models Improve Compositional Tool Calling
Recurrent computation helped most when tool use required multiple dependent calls, not just single API selection.
The paper tests looped language models on API-Bank, BFCL, and NESTful under matched fine-tuning setups. It finds that accuracy on multi-step, dependency-aware tool calling generally improves as recurrent depth increases at inference time. Gains are smaller and more model-dependent for isolated API invocation. Adaptive inference performed better on compute trade-offs by spending extra computation only where needed. HF Daily Papers' note
The paper tests looped language models on API-Bank, BFCL, and NESTful under matched fine-tuning setups. It finds that accuracy on multi-step, dependency-aware tool calling generally improves as recurrent depth increases at inference time. Gains are smaller and more model-dependent for isolated API invocation. Adaptive inference performed better on compute trade-offs by spending extra computation only where needed. HF Daily Papers' note
score 4