Megadose AI progress, ranked and analyzed.

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

· HF Daily Papers ·
Most tested models still missed the benchmark’s health-reasoning questions by a wide margin.

WearableQA uses wearable time series, blood biomarkers, and demographics from 200 real users, with up to 500 days of daily data per person. The benchmark has 4,084 ten-choice questions across 16 question types, separating raw data reasoning from physiological interpretation and single-signal from cross-signal tasks. In tests of 14 proprietary and open-source LLMs, scores ranged from 19.6% to 72.9%, with most models below 60%. HF Daily Papers' note

score 5

Categories: Research