Enhancing VLA Precision in Robotic Manipulation Via FiLM-Based Force/Torque-Vision Integration
We propose a multimodal integration framework to enhance the precision of Vision-Language-Action (VLA) models in contact-rich robotic tasks. Although visual perception is essential for task grounding, it often lacks the force awareness required for high-precision alignment and insertion. To address …