When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback
Mandatory verification restored corrupted-agent math accuracy from 72.4% to 100%.
The paper tests what happens when mathematical agents receive plausible but wrong tool outputs through a hidden interceptor. Across 31 problems, agents without verification fell sharply from perfect accuracy. Same-context mandatory reflection recovered performance completely, while optional verification helped only when models chose to use it. A recovery experiment also found that restarting the full problem after explicit detection succeeded in every case. ArXiv · AI/CL/LG's note
The paper tests what happens when mathematical agents receive plausible but wrong tool outputs through a hidden interceptor. Across 31 problems, agents without verification fell sharply from perfect accuracy. Same-context mandatory reflection recovered performance completely, while optional verification helped only when models chose to use it. A recovery experiment also found that restarting the full problem after explicit detection succeeded in every case. ArXiv · AI/CL/LG's note
score 5