2025
VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation
ICLR 2025poster
Vision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human in…