Megadose AI progress, ranked and analyzed.

Stealing Reasoning Traces from Proprietary LLM APIs

· ArXiv · AI/CL/LG ·
Encrypted reasoning blocks can be replayed into weaker models to make them reveal hidden chain-of-thought in plaintext.

The paper says this works because the encrypted traces are interchangeable across sessions, users, and models inside a provider’s ecosystem. The authors report demonstrations across Anthropic, OpenAI, and Google, framing it as a way around anti-distillation protections. They also say they decoded 315,320 public reasoning blocks and found 367 PII artifacts and 182 credentials. The claimed risks extend to hidden hazardous content and prompt injections carried inside encrypted blocks. ArXiv · AI/CL/LG's note

score 6

Categories: Research