Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
The paper localizes training credit to specific reasoning events in multimodal geometry problems.
It introduces “credit-addressable reasoning,” where the units shown during inference also become the places where learning compares alternatives. The authors implement this with Code-CoT, using line-addressable executable code for visual relations, and CE-GRPO, which assigns localized advantages from outcome differences. Across nine geometry benchmarks, CE-GRPO reports 76.04 average accuracy, ahead of Qwen3-VL-8B by 8.09 points and trajectory-level GRPO by 3.43 points. HF Daily Papers' note
It introduces “credit-addressable reasoning,” where the units shown during inference also become the places where learning compares alternatives. The authors implement this with Code-CoT, using line-addressable executable code for visual relations, and CE-GRPO, which assigns localized advantages from outcome differences. Across nine geometry benchmarks, CE-GRPO reports 76.04 average accuracy, ahead of Qwen3-VL-8B by 8.09 points and trajectory-level GRPO by 3.43 points. HF Daily Papers' note
score 5