Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
The paper says LLM agents can run the code correctly and still reach invalid statistical conclusions.
The authors introduce P-Bench, 425 realistic hypothesis-testing tasks across economics, biology, and medicine. Each task asks an agent to choose a method, compute a p-value, and decide whether the hypothesis holds. Their Fisher-R1 agent is trained with synthetic tasks and reinforcement learning using verified statistical rewards. On P-Bench, Fisher-R1-14B beats its backbone and listed proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro. ArXiv · AI/CL/LG's note
The authors introduce P-Bench, 425 realistic hypothesis-testing tasks across economics, biology, and medicine. Each task asks an agent to choose a method, compute a p-value, and decide whether the hypothesis holds. Their Fisher-R1 agent is trained with synthetic tasks and reinforcement learning using verified statistical rewards. On P-Bench, Fisher-R1-14B beats its backbone and listed proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro. ArXiv · AI/CL/LG's note
score 5