WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Most tested models still missed the benchmark’s health-reasoning questions by a wide margin.
WearableQA uses wearable time series, blood biomarkers, and demographics from 200 real users, with up to 500 days of daily data per person. The benchmark has 4,084 ten-choice questions across 16 question types, separating raw data reasoning from physiological interpretation and single-signal from cross-signal tasks. In tests of 14 proprietary and open-source LLMs, scores ranged from 19.6% to 72.9%, with most models below 60%. HF Daily Papers' note
WearableQA uses wearable time series, blood biomarkers, and demographics from 200 real users, with up to 500 days of daily data per person. The benchmark has 4,084 ten-choice questions across 16 question types, separating raw data reasoning from physiological interpretation and single-signal from cross-signal tasks. In tests of 14 proprietary and open-source LLMs, scores ranged from 19.6% to 72.9%, with most models below 60%. HF Daily Papers' note
score 5