Megadose AI progress, ranked and analyzed.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

· HF Daily Papers ·
The benchmark’s top model solved only 57.3% of scattered mobile-assistant tasks.

SPIEval tests LLMs on 250 human-curated tasks across 4,335 personal records in 10 apps, with multi-turn tool use. The paper frames the work around reasoning, disambiguation, integration, preference inference, and multi-intent decomposition. In the evaluation of nine models, failures were mostly about finding the right information: 79% came from inaccurate localization. The authors also report that advanced search methods appeared in fewer than 2% of retrieval actions. HF Daily Papers' note

score 5

Categories: Research