Megadose AI progress, ranked and analyzed.

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

· ArXiv · AI/CL/LG ·
VAKRA tests whether agents can combine API calls, retrieval, and policy constraints, and the strongest models still break down sharply.

The benchmark covers more than 8,000 executable APIs across 62 domains. Its tasks move from single-hop API use to multi-hop structured reasoning and questions constrained by natural-language tool-use policies. In a fixed ReAct setup, the best model reached 70.4% on single-hop endpoint tasks but fell to about 50-51% on compositional APIs. The paper says failures cluster around entity disambiguation and cross-source grounding, not basic tool-calling mechanics. ArXiv · AI/CL/LG's note

score 6

Categories: Research