← Search

Hyolim Kang

9 accepted papers

2025

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

ICCV 2025poster

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained temporal annotations, we demonstrate that LLMs can effectively group atomic actio…

Cited by 0SourcePDFScholar
2025

UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations

CoRL 2025poster

Mimicry is a fundamental learning mechanism in humans, enabling individuals to learn new tasks by observing and imitating experts. However, applying this ability to robots presents significant challenges due to the inherent differences between human and robot embodiments in both their visual appeara…

Cited by 0SourceScholar
2024

ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos

ECCV 2024poster

"Online Temporal Action Localization (On-TAL) is a critical task that aims to instantaneously identify action instances in untrimmed streaming videos as soon as an action concludes—a major leap from frame-based Online Action Detection (OAD). Yet, the challenge of detecting overlapping actions is oft…

2023

MiniROAD: Minimal RNN Framework for Online Action Detection

ICCV 2023poster

Online Action Detection (OAD) is the task of identifying actions in streaming videos without access to future frames. Much effort has been devoted to effectively capturing long-range dependencies, with transformers receiving the spotlight for their ability to capture long-range temporal structures.…

Cited by 22PDFcodeScholar
2023

Soft-Landing Strategy for Alleviating the Task Discrepancy Problem in Temporal Action Localization Tasks

CVPR 2023poster

Temporal Action Localization (TAL) methods typically operate on top of feature sequences from a frozen snippet encoder that is pretrained with the Trimmed Action Classification (TAC) tasks, resulting in a task discrepancy problem. While existing TAL methods mitigate this issue either by retraining t…

2022

ComMU: Dataset for Combinatorial Music Generation

NeurIPS 2022accept

Commercial adoption of automatic music composition requires the capability of generating diverse and high-quality music suitable for the desired context (e.g., music for romantic movies, action games, restaurants, etc.). In this paper, we introduce combinatorial music generation, a new task to creat…

2022

UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection

CVPR 2022poster

Generic Event Boundary Detection (GEBD) is a newly suggested video understanding task that aims to find one level deeper semantic boundaries of events. Bridging the gap between natural human perception and video understanding, it has various potential applications, including interpretable and semant…

Cited by 36PDFcodeScholar
2021

CAG-QIL: Context-Aware Actionness Grouping via Q Imitation Learning for Online Temporal Action Localization

ICCV 2021poster

Temporal action localization has been one of the most popular tasks in video understanding, due to the importance of detecting action instances in videos. However, not much progress has been made on extending it to work in an online fashion, although many video related tasks can benefit by going onl…

Cited by 14PDFScholar