Megadose AI progress, ranked and analyzed.

Lagged Coupling: Internal Representations Become Readable Before They Become Causal

· ArXiv · AI/CL/LG ·
Probe accuracy showed up early across Pythia, but steering mostly did not move behavior.

The paper reports that target variables were linearly readable from the residual stream by step 1,000 at every tested model scale. Steering along those same probe directions was null-equivalent in 43 of 48 model-checkpoint cells. The author separates the result into internal readability, behavioral readability, and causal efficacy, arguing that representation formation outpaced usable causal readout. An OLMo-2 replication preserved the direction of the finding at weaker magnitude. ArXiv · AI/CL/LG's note

score 5

Categories: Research