Megadose AI progress, ranked and analyzed.

User Model Extraction via Belief Self-Distillation

· ArXiv · AI/CL/LG ·
The paper says an LLM’s hidden belief about the user can be read out and written back to change its behavior.

The authors introduce Belief Self-Distillation, using the frozen model to distill user beliefs from conversations without outside labels. They report that the resulting user representation works across multiple model families and drives stronger interventions than comparable hidden-state steering. In their tests, refusals changed when the inferred user intent was altered while the request stayed the same. They also claim separately trained models show a shared geometry for representing users. ArXiv · AI/CL/LG's note

score 4

Categories: Research