Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
WearableQA:面向真实可穿戴数据健康推理的基准
AI 导读
研究者推出 WearableQA,一个基于 200 名真实用户可穿戴时序、血液生物标志物与人口统计数据的基准,含 4,084 道 10 选项多选题,用户每日测量最长覆盖 500 天。该基准设 16 种题型,从数据与健康推理、单信号与跨信号推理两个维度评估模型能力。14 个闭源与开源 LLM 的准确率在 19.6% 至 72.9% 之间(随机基线 10%),多数模型低于 60%。
HuggingFace Daily Papers(社区热门论文)
39
AI 编辑部评分,满分 100WearableQA:面向真实可穿戴数据健康推理的基准
研究者推出 WearableQA,一个基于 200 名真实用户可穿戴时序、血液生物标志物与人口统计数据的基准,含 4,084 道 10 选项多选题,用户每日测量最长覆盖 500 天。该基准设 16 种题型,从数据与健康推理、单信号与跨信号推理两个维度评估模型能力。14 个闭源与开源 LLM 的准确率在 19.6% 至 72.9% 之间(随机基线 10%),多数模型低于 60%。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org