Structured but Silent: Probing Capability Requirements in LLM Hidden States
The paper finds that LLMs encode tool-capability needs before speaking, but often fail to state them reliably.
The authors test whether hidden states can reveal what kind of external capability a query requires before the model generates an answer. Their TACIT framework splits those requirements across Source, Transformation, and World Effect, producing eight capability classes. Linear probes trained on pre-generation hidden states from four open-weight model families decode those classes with high accuracy. The same models perform worse when asked to classify the queries explicitly in natural language, which the paper calls a representation-to-verbalization gap. ArXiv · AI/CL/LG's note
The authors test whether hidden states can reveal what kind of external capability a query requires before the model generates an answer. Their TACIT framework splits those requirements across Source, Transformation, and World Effect, producing eight capability classes. Linear probes trained on pre-generation hidden states from four open-weight model families decode those classes with high accuracy. The same models perform worse when asked to classify the queries explicitly in natural language, which the paper calls a representation-to-verbalization gap. ArXiv · AI/CL/LG's note
score 5