Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
The paper finds a model can encode a partner’s expertise early, but not use it behaviorally until much later.
Okamoto and Sarti test this on ExpertCollab, a synthetic set of multi-turn research-planning dialogues with model-played personas at four expertise levels. Partner expertise is most decodable in early transformer layers, then drops near chance before the network midpoint. Patching suggests the signal only propagates strongly when inserted past that midpoint, separating representation from causal use. The authors frame this as an initial demonstration with one model and one synthetic corpus. HF Daily Papers' note
Okamoto and Sarti test this on ExpertCollab, a synthetic set of multi-turn research-planning dialogues with model-played personas at four expertise levels. Partner expertise is most decodable in early transformer layers, then drops near chance before the network midpoint. Patching suggests the signal only propagates strongly when inserted past that midpoint, separating representation from causal use. The authors frame this as an initial demonstration with one model and one synthetic corpus. HF Daily Papers' note
score 4