SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
SLCA-GRPO trains tool-calling agents by separating credit for tool decisions from credit for final summaries.
The paper says standard GRPO applies one trajectory-level advantage across all output tokens, letting summary-generation noise affect tool-call tokens. SLCA-GRPO instead assigns execution rewards to tool segments and preference rewards to summary segments within the same rollout group. The authors report better convergence and gains over GRPO, ToolPO, and RLTR on a 7B backbone, including BFCL and tau^2-Bench results. HF Daily Papers' note
The paper says standard GRPO applies one trajectory-level advantage across all output tokens, letting summary-generation noise affect tool-call tokens. SLCA-GRPO instead assigns execution rewards to tool segments and preference rewards to summary segments within the same rollout group. The authors report better convergence and gains over GRPO, ToolPO, and RLTR on a 7B backbone, including BFCL and tau^2-Bench results. HF Daily Papers' note
score 5