Megadose AI progress, ranked and analyzed.

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

· ArXiv · AI/CL/LG ·
PoTRE reports a 49.92% state-of-the-art score on Humanity’s Last Exam using a four-agent reasoning setup.

The paper proposes splitting test-time inference across adversarial refinement, hierarchical planning, spectrum search, and direct chain agents.
A task-adaptive aggregation layer then selects, synthesizes, or verifies the final answer.
The authors say the approach improves reasoning on ARC-AGI-2, HLE, and PRBench Finance while using similar or fewer inference tokens than scaled homogeneous baselines.
Accepted at TMLR 2026.
ArXiv · AI/CL/LG's note

score 4

Categories: Research