Executing Causal Structure Learning with Linear-Attention Transformers
The paper builds a fixed-weight linear-attention transformer that exactly performs one update step of a causal discovery optimizer.
Stacking those blocks reproduces the optimizer’s trajectory, with the current graph and multiplier carried forward between updates. The authors argue the multiplier is necessary because changing it can change the next update. Their tests match a reference update to floating-point precision, but the replayed solver keeps the reference method’s wins and failures on synthetic and benchmark graph topologies. Trained ordinary attention models did not reliably learn the update or generalize to larger graphs under the reported budgets. ArXiv · AI/CL/LG's note
Stacking those blocks reproduces the optimizer’s trajectory, with the current graph and multiplier carried forward between updates. The authors argue the multiplier is necessary because changing it can change the next update. Their tests match a reference update to floating-point precision, but the replayed solver keeps the reference method’s wins and failures on synthetic and benchmark graph topologies. Trained ordinary attention models did not reliably learn the update or generalize to larger graphs under the reported budgets. ArXiv · AI/CL/LG's note
score 4