Model Hypnosis: Strong control of AI via additive subliminal effects
Small prompt cues can add up into strong control over model behavior.
Boix-Adsera and Tessler describe “model hypnosis” as a way seemingly weak, irrelevant prompt details can be combined to steer AI systems. The abstract says the effect appears across model families and sizes, including frontier reasoning models. It also says prompts using these cues can transfer between models. The authors flag inconspicuous choices like paraphrases and typos as safety and interpretability problems. ArXiv · AI/CL/LG's note
Boix-Adsera and Tessler describe “model hypnosis” as a way seemingly weak, irrelevant prompt details can be combined to steer AI systems. The abstract says the effect appears across model families and sizes, including frontier reasoning models. It also says prompts using these cues can transfer between models. The authors flag inconspicuous choices like paraphrases and typos as safety and interpretability problems. ArXiv · AI/CL/LG's note
score 7