VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
The benchmark tests whether agents can keep acting over simulated weeks without being re-prompted.
VibeLifeBench contains 200 long-horizon tasks across ten everyday-life domains, running inside a simulated world with 22 mock services. The world changes on its own clock, including silent changes an agent must discover by checking back. Tasks are graded on the agent’s final state, timing, and handling of implicit constraints. The authors say seven frontier models all scored low, and plan to open-source the tasks, environments, and evaluator. HF Daily Papers' note
VibeLifeBench contains 200 long-horizon tasks across ten everyday-life domains, running inside a simulated world with 22 mock services. The world changes on its own clock, including silent changes an agent must discover by checking back. Tasks are graded on the agent’s final state, timing, and handling of implicit constraints. The authors say seven frontier models all scored low, and plan to open-source the tasks, environments, and evaluator. HF Daily Papers' note
score 5