Megadose AI progress, ranked and analyzed.

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

· HF Daily Papers ·
The paper’s gate improves forecasts by using language models only where past outcomes show they add value beyond an existing baseline.

The authors test a competence-gated pooling method on 2,357 resolved binary questions and five language models. It improves the main external baseline’s Brier score from 0.0771 to 0.0732 and beats global forecast combinations. The gain holds under leakage controls, but not on the official ForecastBench market subset, where the system mostly defers to the market. Verbal confidence from Qwen models was not a reliable signal for when the model should overrule the external forecast. HF Daily Papers' note

score 4

Categories: Research