WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
Eun Chang, Zhuangqun Huang, Yiwei Liao, Sagar Ravi Bhavsar, Amogh Param, Tammy Stark, Adel Ahmadyan, Xiao Yang
Abstract
We introduce WearVQA, the first benchmark specifically designed to evaluate the visual question answering (VQA) capabilities of multi-modal AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on high-quality, third-person imagery, WearVQA reflects the unique chal- lenges of ego-centric interaction—where visual inputs may be occluded, poorly lit, unzoomed, or blurry, and questions are grounded in realistic wearable use cases. The benchmark comprises 2,500 carefully curated image-question-answer triplets, spanning 7 diverse image domains including both text-centric and general scenes, 10 cognitive task types ranging from basic recognition to various forms of reasoning, and 6 common wearables-specific image quality issues. All questions are designed to be answerable using only the visual input and common senses. WearVQA is paired with a rigorous LLM-as-a-judge evaluation framework with 96% labeling accuracy. Open-source and proprietary multi-modal LLMs achieved a QA accuracy as low as 24–52% on WearVQA, with substantial drops on lower-quality images and reasoning- heavy tasks. These observations position WearVQA as a comprehensive and challenging benchmark for guiding technicial advancement towards robust, real-world multi-modal wearables AI systems.
BibTeX
@inproceedings{
chang2025wearvqa,
title={Wear{VQA}: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios},
author={Eun Chang and Zhuangqun Huang and Yiwei Liao and Sagar Ravi Bhavsar and Amogh Param and Tammy Stark and Adel Ahmadyan and Xiao Yang and Jiaqi Wang and Ahsan Abdullah and Giang Nguyen and Akil Iyer and David Patrick hall and Elissa Li and Nicolas SCHEFFER and Ahmed Kirmani and Babak Damavandi and Rakesh Wanga and Anuj Kumar and Rohit Patel and Seungwhan Moon and Xin Luna Dong},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2025},
url={https://openreview.net/forum?id=s5p9ByKN1j}
}