Large Language Models Develop Belief State Geometry In-Context
The paper reports that LLM activations contain a linearly readable belief state when prompted with hidden Markov model data.
Across six open-source models and 40 HMMs, probes recovered posterior hidden-state distributions from residual streams with peak R² values of 0.83 to 0.99. The authors then patched and steered the identified subspace, finding prediction quality stayed near the original model while control interventions degraded it. They argue this is representation-level evidence that in-context learning can approximate Bayesian prediction over a generative model inferred from context. ArXiv · AI/CL/LG's note
Across six open-source models and 40 HMMs, probes recovered posterior hidden-state distributions from residual streams with peak R² values of 0.83 to 0.99. The authors then patched and steered the identified subspace, finding prediction quality stayed near the original model while control interventions degraded it. They argue this is representation-level evidence that in-context learning can approximate Bayesian prediction over a generative model inferred from context. ArXiv · AI/CL/LG's note
score 5