2026
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
ICRA 2026poster
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand–object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Context Fine-Tuning (CCFT)…