2026
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
ICRA 2026poster
Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive datasets infeasible. Existing methods remain limited to closed vocabularies and fixed label sets. In this work, we explore the …