Verifiable Social Reasoning for LLM Assistants
The paper introduces Fuse, a simulation setup meant to give social-advice questions a checkable answer.
Fuse creates interactions where one agent has a hidden motive, then has a user-like agent ask an assistant to infer it. The authors say this gives ground truth “by construction,” and they validate the setup with a human study using 24k annotations. Tested across 12 LLMs, the framework found that user mediation makes the task harder, biased framing affects model judgments, and longer chats do not reliably improve performance. The team is releasing Fuse and a 21k-example dataset. HF Daily Papers' note
Fuse creates interactions where one agent has a hidden motive, then has a user-like agent ask an assistant to infer it. The authors say this gives ground truth “by construction,” and they validate the setup with a human study using 24k annotations. Tested across 12 LLMs, the framework found that user mediation makes the task harder, biased framing affects model judgments, and longer chats do not reliably improve performance. The team is releasing Fuse and a 21k-example dataset. HF Daily Papers' note
score 4