← Search

Minho Shim

11 accepted papers

2026

Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning Segmentation

ICLR 2026poster

Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to instructions. To exploit this for training-free localization, we cast video reasoning segmentation as video QA and extract attention maps via rollout. Since raw maps are too nois…

Cited by 0SourcecodeScholar
2025

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

ICCV 2025poster

Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insigh…

2025

Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval

ICCV 2025poster

In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhance…

Cited by 0SourcePDFScholar
2024

Classification Matters: Improving Video Action Detection with Class-Specific Attention

ECCV 2024oral

"Video action detection (VAD) aims to detect actors and classify their actions in a video. We figure that VAD suffers more from classification rather than localization of actors. Hence, we analyze how prevailing methods form features for classification and find that they prioritize actor regions, ye…

Cited by 0SourcePDFScholar
2024

Towards Multi-Domain Learning for Generalizable Video Anomaly Detection

NeurIPS 2024poster

Most of the existing Video Anomaly Detection (VAD) studies have been conducted within single-domain learning, where training and evaluation are performed on a single dataset. However, the criteria for abnormal events differ across VAD datasets, making it problematic to apply a single-domain model to…

Cited by 1SourcePDFScholar
2023

Decomposed Cross-Modal Distillation for RGB-Based Temporal Action Detection

CVPR 2023poster

Temporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit slow inference speed due to their reliance on computationally expensive optical flow. In this paper, we introduce a dec…

Cited by 22SourcePDFScholar
2023

Exploring Temporally Dynamic Data Augmentation for Video Recognition

ICLR 2023top-25%

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been rarely explored despite its effectiveness. Few existing augmentation recipes for video recognition naively extend the im…

Cited by 13SourcePDFScholar
2023

Frequency Selective Augmentation for Video Representation Learning

AAAI 2023technical

Recent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a mechanism to prevent representation learning from bias toward…

Cited by 5SourcePDFScholar
2020

Learning from Dances: Pose-Invariant Re-Identification for Multi-Person Tracking

ICASSP 2020accepted

Most existing multi-person tracking approaches rely on appearance based re-identification (re-ID) to resolve fragmented tracklets. However, simply using appearance information could be insufficient for videos containing severe pose changes, such as sports or dance videos. With the goal of learning p…

Cited by 0SourceScholar
2020

READ: Reciprocal Attention Discriminator for Image-to-Video Re-Identification

ECCV 2020poster

Person re-identification (re-ID) is the problem of visually identifying a person given a database of identities. In this work, we focus on image-to-video re-ID which compares a single query image to videos in the gallery. The main challenge is the asymmetry association of an image and a video, and o…

2018

Teaching Machines to Understand Baseball Games: Large-Scale Baseball Video Database for Multiple Video Understanding Tasks

ECCV 2018poster

A major obstacle in teaching machines to understand videos is the lack of training data, as creating temporal annotations for long videos requires a huge amount of human effort. To this end, we introduce a new large-scale baseball video dataset called the BBDB, which is produced semi-automatically b…

Cited by 16SourcePDFScholar