VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
VoiceMem splits speech memory into factual and emotional tracks built for live conversations.
The paper describes a dual “left brain” and “right brain” memory system for duplex speech language models. Its factual retriever is reported to beat Mem0’s top-200 result by nearly 30 points while using top-5 retrieval. The affective side models persona and emotion across short and long horizons, improving the aggregate benchmark score by 4.29 points over the prior best system. Retrieval is reported at 134 ms, which the authors say fits inside standard VAD latency without adding conversational delay. ArXiv · AI/CL/LG's note
The paper describes a dual “left brain” and “right brain” memory system for duplex speech language models. Its factual retriever is reported to beat Mem0’s top-200 result by nearly 30 points while using top-5 retrieval. The affective side models persona and emotion across short and long horizons, improving the aggregate benchmark score by 4.29 points over the prior best system. Retrieval is reported at 134 ms, which the authors say fits inside standard VAD latency without adding conversational delay. ArXiv · AI/CL/LG's note
score 4