Megadose AI progress, ranked and analyzed.

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

· ArXiv · AI/CL/LG ·
The benchmark tests LLMs on up to 500 days of real wearable records from 200 users.

WearableQA contains 4,084 ten-choice questions using wearable time series, blood biomarkers, and demographics. The authors split the tasks across data computation, health interpretation, single-signal reasoning, and cross-signal reasoning. In their evaluation of 14 proprietary and open-source models, accuracy ranged from 19.6% to 72.9%, with most models below 60%. ArXiv · AI/CL/LG's note

score 4

Categories: Research