Megadose AI progress, ranked and analyzed.

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

· HF Daily Papers ·
The paper localizes training credit to specific reasoning events in multimodal geometry problems.

It introduces “credit-addressable reasoning,” where the units shown during inference also become the places where learning compares alternatives. The authors implement this with Code-CoT, using line-addressable executable code for visual relations, and CE-GRPO, which assigns localized advantages from outcome differences. Across nine geometry benchmarks, CE-GRPO reports 76.04 average accuracy, ahead of Qwen3-VL-8B by 8.09 points and trajectory-level GRPO by 3.43 points. HF Daily Papers' note

score 5

Categories: Research