← Search

Dongyoon Wee

19 accepted papers

2026

Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning Segmentation

ICLR 2026poster

Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to instructions. To exploit this for training-free localization, we cast video reasoning segmentation as video QA and extract attention maps via rollout. Since raw maps are too nois…

Cited by 0SourcecodeScholar
2026

SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models

CVPR 2026

Diffusion models are a strong backbone for visual generation, but their inherently sequential denoising process leads to slow inference. Previous methods accelerate sampling by caching and reusing intermediate outputs based on feature distances between adjacent timesteps. However, existing caching s

Cited by 0SourcecodeScholar
2025

CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images

CVPR 2025poster

3D Gaussian Splatting (3DGS) has attracted significant attention for its high-quality novel view rendering, inspiring research to address real-world challenges. While conventional methods depend on sharp images for accurate scene reconstruction, real-world scenarios are often affected by defocus blu…

Cited by 0SourcePDFScholar
2025

CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images

ICCV 2025poster

3D Gaussian Splatting (3DGS) has gained significant attention for their high-quality novel view rendering, motivating research to address real-world challenges. A critical issue is the camera motion blur caused by movement during exposure, which hinders accurate 3D scene reconstruction. In this stud…

2025

Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos

ICCV 2025poster

Recent works on dynamic 3D neural field reconstruction assume the input from synchronized multi-view videos whose poses are known. The input constraints are often not satisfied in real-world setups, making the approach impractical. We show that unsynchronized videos from unknown poses can generate d…

Cited by 0SourcePDFScholar
2025

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

ICCV 2025poster

Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insigh…

2025

Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval

ICCV 2025poster

In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhance…

Cited by 0SourcePDFScholar
2024

Classification Matters: Improving Video Action Detection with Class-Specific Attention

ECCV 2024oral

"Video action detection (VAD) aims to detect actors and classify their actions in a video. We figure that VAD suffers more from classification rather than localization of actors. Hence, we analyze how prevailing methods form features for classification and find that they prioritize actor regions, ye…

Cited by 0SourcePDFScholar
2024

Motion-Oriented Compositional Neural Radiance Fields for Monocular Dynamic Human Modeling

ECCV 2024poster

"This paper introduces Motion-oriented Compositional Neu-ral Radiance Fields (MoCo-NeRF), a framework designed to perform free-viewpoint rendering of monocular human videos via novel non-rigid motion modeling approach. In the context of dynamic clothed humans, complex cloth dynamics generate non-rig…

2024

Regularizing Dynamic Radiance Fields with Kinematic Fields

ECCV 2024poster

"This paper presents a novel approach for reconstructing dynamic radiance fields from monocular videos. We integrate kinematics with dynamic radiance fields, bridging the gap between the sparse nature of monocular videos and the real-world physics. Our method introduces the kinematic field, capturin…

Cited by 0SourcePDFScholar
2024

Towards Multi-Domain Learning for Generalizable Video Anomaly Detection

NeurIPS 2024poster

Most of the existing Video Anomaly Detection (VAD) studies have been conducted within single-domain learning, where training and evaluation are performed on a single dataset. However, the criteria for abnormal events differ across VAD datasets, making it problematic to apply a single-domain model to…

Cited by 1SourcePDFScholar
2023

Decomposed Cross-Modal Distillation for RGB-Based Temporal Action Detection

CVPR 2023poster

Temporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit slow inference speed due to their reliance on computationally expensive optical flow. In this paper, we introduce a dec…

Cited by 22SourcePDFScholar
2023

Exploring Temporally Dynamic Data Augmentation for Video Recognition

ICLR 2023top-25%

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been rarely explored despite its effectiveness. Few existing augmentation recipes for video recognition naively extend the im…

Cited by 13SourcePDFScholar
2023

Frequency Selective Augmentation for Video Representation Learning

AAAI 2023technical

Recent self-supervised video representation learning methods focus on maximizing the similarity between multiple augmented views from the same video and largely rely on the quality of generated views. However, most existing methods lack a mechanism to prevent representation learning from bias toward…

Cited by 5SourcePDFScholar
2023

SEFD: Learning to Distill Complex Pose and Occlusion

ICCV 2023poster

This paper addresses the problem of three-dimensional (3D) human mesh estimation in complex poses and occluded situations. Although many improvements have been made in 3D human mesh estimation using the two-dimensional (2D) pose with occlusion between humans, occlusion from complex poses and other o…

Cited by 13PDFcodeScholar
2020

Learning from Dances: Pose-Invariant Re-Identification for Multi-Person Tracking

ICASSP 2020accepted

Most existing multi-person tracking approaches rely on appearance based re-identification (re-ID) to resolve fragmented tracklets. However, simply using appearance information could be insufficient for videos containing severe pose changes, such as sports or dance videos. With the goal of learning p…

Cited by 0SourceScholar
2020

READ: Reciprocal Attention Discriminator for Image-to-Video Re-Identification

ECCV 2020poster

Person re-identification (re-ID) is the problem of visually identifying a person given a database of identities. In this work, we focus on image-to-video re-ID which compares a single query image to videos in the gallery. The main challenge is the asymmetry association of an image and a video, and o…

2020

Regularization on Spatio-Temporally Smoothed Feature for Action Recognition

CVPR 2020poster

Deep neural networks for video action recognition frequently require 3D convolutional filters and often encounter overfitting due to a larger number of parameters. In this paper, we propose Random Mean Scaling (RMS), a simple and effective regularization method, to relieve the overfitting problem in…

Cited by 34PDFScholar