Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
The paper finds that suppressing an “evaluation-awareness” latent from the prompt side does not reliably suppress evaluation-aware behavior.
The authors optimize fluent prompts to drive selected internal features toward zero without inference-time activation edits. On Llama-3.2-3B and Llama-3.1-8B, they can strongly suppress some targeted latents, including fully turning off a validated SAE feature. But the CAA evaluation direction behaves poorly as a control target: a random direction is just as suppressible, and optimizing a prefix around a real eval passage slightly increases the model’s eval judgment. Their conclusion is that activation readability is not the same as behavioral controllability. ArXiv · AI/CL/LG's note
The authors optimize fluent prompts to drive selected internal features toward zero without inference-time activation edits. On Llama-3.2-3B and Llama-3.1-8B, they can strongly suppress some targeted latents, including fully turning off a validated SAE feature. But the CAA evaluation direction behaves poorly as a control target: a random direction is just as suppressible, and optimizing a prefix around a real eval passage slightly increases the model’s eval judgment. Their conclusion is that activation readability is not the same as behavioral controllability. ArXiv · AI/CL/LG's note
score 5