AAAI 2026technical0 citations
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
Wenxuan Song, Ziyang Zhou, Han Zhao, Jiayi Chen, Pengxiang Ding, Haodong Yan, Yuxin Huang, Feilong Tang
Abstract
Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispersed. To guide the visual attention grounding on the correct target, we propose ReconVLA, a reconstructive VLA model with an implicit grounding paradigm. Conditioned on the model
BibTeX
@inproceedings{aaai2026_reconvlareconstr,
title = {ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver},
author = {Wenxuan Song and Ziyang Zhou and Han Zhao and Jiayi Chen and Pengxiang Ding and Haodong Yan and Yuxin Huang and Feilong Tang and Donglin Wang and Haoang Li},
booktitle = {AAAI 2026},
year = {2026}
}