Megadose AI progress, ranked and analyzed.

SocietyBench: Forecasting Counterfactual Social-World Evolution

· ArXiv · AI/CL/LG ·
SocietyBench tests whether LLMs can forecast anonymized social-event timelines, and the best model still scores only 75/100.

The benchmark turns real news and social-media arcs into counterfactual versions by replacing named entities and shifting dates, so models cannot lean on memorized events. It scores answers separately for probability calibration and temporal accuracy, which the paper says can diverge. Six frontier LLMs were tested across five events and 125 prediction points in Chinese and English. Agent frameworks did not improve on their shared base model, and the authors released the anonymized timelines, questions, ground truth, and scoring code. ArXiv · AI/CL/LG's note

score 4

Categories: Research