DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
The benchmark’s strongest baseline stays under 45% on dense, stateful visual-memory tasks.
DSV-Mem tests MLLM agents on professional workflows where structured artifacts change over time.
It includes 1,000 expert-reviewed questions across current state, past state, derived state, change history, and conflict/refusal cases.
The authors find that state evolution, especially governing updates, is the main difficulty, more than OCR, arithmetic, or raw context length.
They report limited gains from extra reasoning effort and memory-management methods, while state-aware designs perform better.
ArXiv · AI/CL/LG's note
DSV-Mem tests MLLM agents on professional workflows where structured artifacts change over time.
It includes 1,000 expert-reviewed questions across current state, past state, derived state, change history, and conflict/refusal cases.
The authors find that state evolution, especially governing updates, is the main difficulty, more than OCR, arithmetic, or raw context length.
They report limited gains from extra reasoning effort and memory-management methods, while state-aware designs perform better.
ArXiv · AI/CL/LG's note
score 5