Megadose AI progress, ranked and analyzed.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

· ArXiv · AI/CL/LG ·
The paper says LLM agents can run the code correctly and still reach invalid statistical conclusions.

The authors introduce P-Bench, 425 realistic hypothesis-testing tasks across economics, biology, and medicine. Each task asks an agent to choose a method, compute a p-value, and decide whether the hypothesis holds. Their Fisher-R1 agent is trained with synthetic tasks and reinforcement learning using verified statistical rewards. On P-Bench, Fisher-R1-14B beats its backbone and listed proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro. ArXiv · AI/CL/LG's note

score 5

Categories: Research