RUMBA: Russian User Memory Benchmark
RUMBA tests whether memory systems can reason over Russian, timestamped conversations across sessions.
The benchmark pairs user-assistant dialogues with questions that require retrieval, combination, and temporal reasoning. Its taxonomy separates memory tasks by semantic type, session scope, and how explicit the time references are. The authors also include an aligned English subset, but the benchmark is designed around Russian. They present it as a diagnostic for comparing long-context models and memory mechanisms across specific failure modes. ArXiv · AI/CL/LG's note
The benchmark pairs user-assistant dialogues with questions that require retrieval, combination, and temporal reasoning. Its taxonomy separates memory tasks by semantic type, session scope, and how explicit the time references are. The authors also include an aligned English subset, but the benchmark is designed around Russian. They present it as a diagnostic for comparing long-context models and memory mechanisms across specific failure modes. ArXiv · AI/CL/LG's note
score 4