← Search

Tiantian Zheng

2 accepted papers

2026

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos

ICRA 2026poster

Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand–object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Context Fine-Tuning (CCFT)…

Cited by 0codeScholar
2026

Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning

CVPR 2026

Dual-hand action segmentation, densely predicting actions for both hands from untrimmed videos, is essential for understanding complex bimanual activities. However, it poses several unique challenges: complex inter-hand dependencies, visual asymmetry between hands, representation conflicts where the

Cited by 0SourcecodeScholar