Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
The paper finds that causal knowledge may be in the model, but observational context can flip how it is used.
In controlled Simpson’s-paradox settings, adding more interventional samples during pretraining did not by itself make models choose the correct causal direction. The decisive factor was the evidence type shown at inference: observational contexts often induced sign reversals, while aligned interventional probes performed much better. Removing observational evidence released a suppressed causal interpolation ability, and activation patching localized the switch to middle-layer observational rows. ArXiv · AI/CL/LG's note
In controlled Simpson’s-paradox settings, adding more interventional samples during pretraining did not by itself make models choose the correct causal direction. The decisive factor was the evidence type shown at inference: observational contexts often induced sign reversals, while aligned interventional probes performed much better. Removing observational evidence released a suppressed causal interpolation ability, and activation patching localized the switch to middle-layer observational rows. ArXiv · AI/CL/LG's note
score 4