SanSi: A Looped Typed Decision Model for System 1.5 Thinking
SanSi tests whether repeated hidden-state loops can buy reasoning depth without token-by-token explanations.
The paper turns a pre-trained looped language model into a typed decision model, reading option probabilities after each loop and training every loop with a proper scoring rule. On 10,027 test decisions from 59 sources, it reports 72.0% accuracy, well above same-shape and newer non-looped baselines. The authors say loops also generalize to deeper controlled tasks beyond training depth, and that SanSi can serve as a reinforcement-learning judge without gold answers. Source: HF Daily Papers' note
The paper turns a pre-trained looped language model into a typed decision model, reading option probabilities after each loop and training every loop with a proper scoring rule. On 10,027 test decisions from 59 sources, it reports 72.0% accuracy, well above same-shape and newer non-looped baselines. The authors say loops also generalize to deeper controlled tasks beyond training depth, and that SanSi can serve as a reinforcement-learning judge without gold answers. Source: HF Daily Papers' note
score 4