CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
The paper argues that causal discovery benchmarks can mislead when pretrained models have seen similar synthetic worlds.
CausalArena standardizes evaluation across synthetic, semantic operational, formula-grounded, and real-world SCM settings. The authors report that classical, neural, and pretrained methods change rank substantially across benchmark families and protocols. Their conclusion is that benchmark diversity and pretraining-test overlap have become core evaluation problems for causal discovery models. ArXiv · AI/CL/LG's note
CausalArena standardizes evaluation across synthetic, semantic operational, formula-grounded, and real-world SCM settings. The authors report that classical, neural, and pretrained methods change rank substantially across benchmark families and protocols. Their conclusion is that benchmark diversity and pretraining-test overlap have become core evaluation problems for causal discovery models. ArXiv · AI/CL/LG's note
score 5