← Search

Zhang Min

1 accepted papers

2025

VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation

ICLR 2025poster

Vision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human in…