Megadose Built for builders and researchers.

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

· HF Daily Papers ·
RealCompanion tests long-memory reasoning on ten real AI companion relationships, not synthetic user histories.

The benchmark covers 27,218 messages across conversations lasting up to 120 days. Its materials include derived profiles, personas, chat ground truth, and question sets, each tied back to cited messages. The paper reports that memory is rarely needed, and when it is, the relevant message is often far back. It also says tested detectors failed to reliably tell when memory was needed, while agent systems reached similar persona-reconstruction F1 at sharply different costs. HF Daily Papers' note

score 5

Categories: Research