ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
ParaTempo uses “temporal confidence” to cut wasted parallel reasoning without training a new model.
The framework periodically probes each reasoning branch for its tentative answer distribution, then tracks whether recent probes are converging on one answer. Branches with low confidence can be pruned, stable branches can be retired early, and freed compute can be used to fork new branches. In the paper’s math and science benchmarks, it reduced average latency by 21.8-32.2% and token use by 18.1-30.3% while keeping accuracy competitive. HF Daily Papers' note
The framework periodically probes each reasoning branch for its tentative answer distribution, then tracks whether recent probes are converging on one answer. Branches with low confidence can be pruned, stable branches can be retired early, and freed compute can be used to fork new branches. In the paper’s math and science benchmarks, it reduced average latency by 21.8-32.2% and token use by 18.1-30.3% while keeping accuracy competitive. HF Daily Papers' note
score 5