When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
Representation engineering helps most when data or compute is tight, but it does not beat behavioral safeguards overall.
The paper compares behavioral safety methods with approaches that read or steer internal model representations under matched settings. DPO gives the strongest safety control overall, though its gains can erode after later benign fine-tuning. Representation steering is most competitive in low-data cases, especially with strong contrastive data. For monitoring, specialized text monitors are more accurate, while representation probes offer lower marginal cost and can help recover safety lost after fine-tuning. HF Daily Papers' note
The paper compares behavioral safety methods with approaches that read or steer internal model representations under matched settings. DPO gives the strongest safety control overall, though its gains can erode after later benign fine-tuning. Representation steering is most competitive in low-data cases, especially with strong contrastive data. For monitoring, specialized text monitors are more accurate, while representation probes offer lower marginal cost and can help recover safety lost after fine-tuning. HF Daily Papers' note
score 4