← Search

Kyle Min

11 accepted papers

2025

ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning

ICCV 2025poster

In this work, we tackle the problem of video class-incremental learning (VCIL). Many existing VCIL methods mitigate catastrophic forgetting by rehearsal training with a few temporally dense samples stored in episodic memory, which is memory-inefficient. Alternatively, some methods store temporally s…

Cited by 0SourcePDFScholar
2025

Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions

NeurIPS 2025poster

Despite recent advances, vision-language models trained with standard contrastive objectives still struggle with compositional reasoning -- the ability to understand structured relationships between visual and linguistic elements. This shortcoming is largely due to the tendency of the text encoder t…

Cited by 0SourceScholar
2025

EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment

NeurIPS 2025spotlight

Erasing harmful or proprietary concepts from powerful text‑to‑image generators is an emerging safety requirement, yet current ``concept erasure'' techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myo…

Cited by 0SourceScholar
2024

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

CVPR 2024poster

We present Egocentric Action Scene Graphs (EASGs) a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos such as verb-noun action labels by providing a temporally evolving graph-based description of the act…

2024

WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

CVPR 2024poster

The rapid advancement of generative models facilitating the creation of hyper-realistic images from textual descriptions has concurrently escalated critical societal concerns such as misinformation. Although providing some mitigation traditional fingerprinting mechanisms fall short in attributing re…

2023

SViTT: Temporal Learning of Sparse Video-Text Transformers

CVPR 2023poster

Do video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has revealed the strong tendency of video-text models towards frame-based spatial representations, while temporal reasoning remain…

2022

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

ECCV 2022poster

"Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a novel spatial-temporal graph learning framework that can solv…

2020

Adversarial Background-Aware Loss for Weakly-supervised Temporal Activity Localization

ECCV 2020poster

Temporally localizing activities within untrimmed videos has been extensively studied in recent years. Despite recent advances, existing methods for weakly-supervised temporal activity localization struggle to recognize when an activity is not occurring. To address this issue, we propose a novel met…

2019

TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection

ICCV 2019poster

TASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several consecutive frames, and then the following prediction network decodes…

Cited by 209PDFcodeScholar
2018

Hierarchical Novelty Detection for Visual Object Recognition

CVPR 2018poster

Deep neural networks have achieved impressive success in large-scale visual object recognition tasks with a predefined set of classes. However, recognizing objects of novel classes unseen during training still remains challenging. The problem of detecting such novel classes has been addressed in the…

Cited by 94SourcePDFScholar