ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
ToolSciVer adds visual tools to help VLMs verify claims against figures, tables, charts, and surrounding text in scientific papers.
The framework gives a vision-language model table focus, chart parsing, and high-resolution zoom tools so dense visuals can be turned into explicit evidence. Its policy is trained with GRPO using rewards for correctness, formatting, length control, efficient tool use, and valid tool calls. The authors report stronger results than four prompting- and RL-based baselines on SciVer and MuSciClaims across five VLMs from Qwen, InternVL, and Gemma. ArXiv · AI/CL/LG's note
The framework gives a vision-language model table focus, chart parsing, and high-resolution zoom tools so dense visuals can be turned into explicit evidence. Its policy is trained with GRPO using rewards for correctness, formatting, length control, efficient tool use, and valid tool calls. The authors report stronger results than four prompting- and RL-based baselines on SciVer and MuSciClaims across five VLMs from Qwen, InternVL, and Gemma. ArXiv · AI/CL/LG's note
score 5