2026
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
RSS 2026poster
Vision–Language–Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most existing VLAs rely primarily on 2D visual representations, which limits their ability to reason about fine-grained geometry…