Enhancing VLA Precision in Robotic Manipulation Via FiLM-Based Force/Torque-Vision Integration
Abstract
We propose a multimodal integration framework to enhance the precision of Vision-Language-Action (VLA) models in contact-rich robotic tasks. Although visual perception is essential for task grounding, it often lacks the force awareness required for high-precision alignment and insertion. To address this limitation, we leverage Feature-wise Linear Modulation (FiLM) to condition intermediate visual representations on 6-axis Force/Torque (F/T) data. This lightweight fusion strategy allows the model to modulate its action predictions based on real-time physical resistance without incurring significant computational overhead. Experimental results on a UR5e manipulator demonstrate that the proposed F/T-Vision integration enhances contact stability and precision in demanding manipulation tasks compared with vision-only baselines.