The First Token Is a Clue: Verbalizing Multi-Token Concepts from the J-lens
J-lens can recover multi-token concepts by using the first token as the handle.
The paper says the first token of a multi-token concept is nearly as readable as a single-token concept. With that first token and the source prompt, the frozen model recovered the second token in 88.3% of two-token cases. The authors then recover full concept vectors from later hidden states and score them against the vocabulary. Across tests on Gemma, Llama, and Qwen models, their method beat Template Lens on Rank@10 and causal concept swaps. ArXiv · AI/CL/LG's note
The paper says the first token of a multi-token concept is nearly as readable as a single-token concept. With that first token and the source prompt, the frozen model recovered the second token in 88.3% of two-token cases. The authors then recover full concept vectors from later hidden states and score them against the vocabulary. Across tests on Gemma, Llama, and Qwen models, their method beat Template Lens on Rank@10 and causal concept swaps. ArXiv · AI/CL/LG's note
score 4