Megadose AI progress, ranked and analyzed.

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

· ArXiv · AI/CL/LG ·
TRACE turns causal diagnosis into an RL task by hiding simulator interventions as verifiable labels.

The paper builds a digital-advertising diagnostic environment where agents use Python and SQL to find one of 12 root causes and, when needed, the affected segment. Its synthesized rewards come from interventions injected into a controlled simulator, avoiding the need for expert-confirmed real-world causes. On 235 held-out episodes, RL-trained Qwen3.5-35B-A3B reaches 0.757 FullAttr@1, above the strongest prompted baseline cited, Claude Opus 5 at 0.686. The authors argue the result points to scalable objective training signals as a bigger constraint than model scale alone. ArXiv · AI/CL/LG's note

score 5

Categories: Research