When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
Fine-tuning made the oracle selectively worse at naming the very hidden concept it was trained around.
In a controlled Taboo Word Guessing setup, the subject model used a hidden concept while avoiding direct disclosure. The paper finds that fine-tuned activation oracles could become “anti-readers,” persistently failing to recover that training concept. The target was still decodable inside the oracle, pointing to a failure in the oracle’s readout path rather than simple absence of the information. HF Daily Papers' note
In a controlled Taboo Word Guessing setup, the subject model used a hidden concept while avoiding direct disclosure. The paper finds that fine-tuned activation oracles could become “anti-readers,” persistently failing to recover that training concept. The target was still decodable inside the oracle, pointing to a failure in the oracle’s readout path rather than simple absence of the information. HF Daily Papers' note
score 5