Toward SLM-based agentic task-tool intent matching
The paper tests small language models as per-call gatekeepers for whether an agent’s tool choice fits its assigned task.
The authors argue that normal authorization can say a tool call is allowed, but not whether it is a relevant step toward the task. Their setup treats an SLM as a task-tool relevance classifier, judging each selected tool independently before downstream enforcement. They use a new dataset of multi-tool tasks spanning distinct MCP servers, then try prompt optimization, supervised fine-tuning, and GRPO to specialize the models. ArXiv · AI/CL/LG's note
The authors argue that normal authorization can say a tool call is allowed, but not whether it is a relevant step toward the task. Their setup treats an SLM as a task-tool relevance classifier, judging each selected tool independently before downstream enforcement. They use a new dataset of multi-tool tasks spanning distinct MCP servers, then try prompt optimization, supervised fine-tuning, and GRPO to specialize the models. ArXiv · AI/CL/LG's note
score 4