Megadose AI progress, ranked and analyzed.

Stealing Reasoning Traces from Proprietary LLM APIs

Simon Willison ·
Encrypted reasoning blocks could be replayed into weaker sibling models and exposed as plaintext.

Simon Willison points to a paper showing that Anthropic, OpenAI, and Google returned encrypted chain-of-thought blocks that could be reused across sessions, users, and models. The authors found same-family models shared encryption keys, letting attackers feed traces from stronger models into weaker ones and jailbreak them into revealing hidden reasoning. Willison notes the providers acknowledged the reports and the same attacks no longer worked afterward. Simon Willison's note

score 6

Categories: Research