ToolGrad: Efficient tool-use dataset generation with textual "gradients"
Google says generating the tool-use chain first made synthetic agent-training data cheaper, more reliable, and harder.
ToolGrad builds verified API workflows before writing the matching user prompt, reversing the query-first approach used in earlier dataset generation. Google reports a 99.8% pass rate, lower generation cost, and more complex long-horizon samples. Gemma-3 models fine-tuned on a 500-sample ToolGrad dataset improved across tested sizes, with ToolGrad-12B scoring 83.1 on BFCL, near Gemini 2.5 Pro’s 83.2 and above the proprietary models Google lists except that mark. The work was presented at ACL 2026, with future work aimed at broader API ecosystems and on-the-fly personalization. Google Research's note
ToolGrad builds verified API workflows before writing the matching user prompt, reversing the query-first approach used in earlier dataset generation. Google reports a 99.8% pass rate, lower generation cost, and more complex long-horizon samples. Gemma-3 models fine-tuned on a 500-sample ToolGrad dataset improved across tested sizes, with ToolGrad-12B scoring 83.1 on BFCL, near Gemini 2.5 Pro’s 83.2 and above the proprietary models Google lists except that mark. The work was presented at ACL 2026, with future work aimed at broader API ecosystems and on-the-fly personalization. Google Research's note
score 5