MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
The benchmark tests whether mobile agents can plan, call tools, use memory, and recover from runtime constraints in a live sandbox.
MobilePA-Bench spans 13 functional domains and 212 realistic mobile tools, with live app databases and structured feedback. It evaluates sub-agent collaboration, memory use, and pre-packaged skill use, beyond simple GUI manipulation or static API matching. The paper reports that frontier LLMs still perform unreliably when tool order, permissions, and unexpected errors matter. HF Daily Papers' note
MobilePA-Bench spans 13 functional domains and 212 realistic mobile tools, with live app databases and structured feedback. It evaluates sub-agent collaboration, memory use, and pre-packaged skill use, beyond simple GUI manipulation or static API matching. The paper reports that frontier LLMs still perform unreliably when tool order, permissions, and unexpected errors matter. HF Daily Papers' note
score 5