← Search

Kyungho Bae

3 accepted papers

2025

ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning

ICCV 2025poster

In this work, we tackle the problem of video class-incremental learning (VCIL). Many existing VCIL methods mitigate catastrophic forgetting by rehearsal training with a few temporally dense samples stored in episodic memory, which is memory-inefficient. Alternatively, some methods store temporally s…

Cited by 0SourcePDFScholar
2025

MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations

CVPR 2025highlight

In this work, we tackle action-scene hallucination in Video Large Language Models (Video-LLMs), where models incorrectly predict actions based on the scene context or scenes based on observed actions. We observe that existing Video-LLMs often suffer from action-scene hallucination due to two main fa…

Cited by 2SourcePDFScholar
2024

DEVIAS: Learning Disentangled Video Representations of Action and Scene

ECCV 2024oral

"Video recognition models often learn scene-biased action representation due to the spurious correlation between actions and scenes in the training data. Such models show poor performance when the test data consists of videos with unseen action-scene combinations. Although Scene-debiased action reco…