From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge
The paper finds a staged handoff in how LLMs move from routing a question to relying on answer content.
The authors probe hidden states layer by layer across Qwen, Llama, and Gemma on country-continent and related answer tasks. In Qwen, a country-specific request direction becomes strong before interventions on it start changing later fitted knowledge, while the answer content is still taking shape. The three models do not behave uniformly: Gemma shows partial overlap between routing and content in mid-layers, while Llama lacks a sustained routing-effect window under the same tests. In paired questions, dependence on a global request direction falls in later layers, but fitted content remains influential. ArXiv · AI/CL/LG's note
The authors probe hidden states layer by layer across Qwen, Llama, and Gemma on country-continent and related answer tasks. In Qwen, a country-specific request direction becomes strong before interventions on it start changing later fitted knowledge, while the answer content is still taking shape. The three models do not behave uniformly: Gemma shows partial overlap between routing and content in mid-layers, while Llama lacks a sustained routing-effect window under the same tests. In paired questions, dependence on a global request direction falls in later layers, but fitted content remains influential. ArXiv · AI/CL/LG's note
score 5