← Search

Xiangyi Wei

2 accepted papers

2026

Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation

ICRA 2026poster

The Vision-Language-Action models (VLA) have achieved significant advances in robotic manipulation recently. However, vision-only VLA models create fundamental limitations, particularly in perceiving interactive and manipulation dynamic processes. This paper proposes Audio-VLA, a multimodal manipula…

2026

S²-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation

IJCAI 2026

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulative error propagation. This limitation largely arises from static feature fusion mechanisms that rely on fixed weights t

Cited by 0Scholar