Scaling Inherently Interpretable Language Models
The paper claims interpretability can improve as language models scale, instead of trading off against capability.
The authors describe a training pipeline where interpretability is optimized alongside the language-modeling objective. Across three orders of magnitude of compute, they report more disentangled representations and stronger alignment with human-understandable concepts. Their Steerling-8B diffusion language model can attribute generated tokens to input tokens, concepts, and training data, then use those attributions for concept steering without retraining. HF Daily Papers' note
The authors describe a training pipeline where interpretability is optimized alongside the language-modeling objective. Across three orders of magnitude of compute, they report more disentangled representations and stronger alignment with human-understandable concepts. Their Steerling-8B diffusion language model can attribute generated tokens to input tokens, concepts, and training data, then use those attributions for concept steering without retraining. HF Daily Papers' note
score 5