← Search

Jongseo Lee

3 accepted papers

2025

Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition

NeurIPS 2025spotlight

Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Langu…

Cited by 0SourceScholar
2025

ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning

ICCV 2025poster

In this work, we tackle the problem of video class-incremental learning (VCIL). Many existing VCIL methods mitigate catastrophic forgetting by rehearsal training with a few temporally dense samples stored in episodic memory, which is memory-inefficient. Alternatively, some methods store temporally s…

Cited by 0SourcePDFScholar
2023

CAST: Cross-Attention in Space and Time for Video Action Recognition

NeurIPS 2023poster

Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), that achieves a…