← Search

Suyan Dai

1 accepted papers

2026

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

ICLR 2026poster

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image–text–action data and remain limited in two key ways: (i) they struggle with pixel-level scene understanding, and (…

Cited by 0SourceScholar