← Search

Sangyoun Lee

29 accepted papers

2026

MoRGS: Efficient Per-Gaussian Motion Reasoning for Streamable Dynamic 3D Scenes

CVPR 2026

Online reconstruction of dynamic scenes aims to learn from streaming multi-view inputs under low-latency constraints. The fast training and real-time rendering capabilities of 3D Gaussian Splatting have made on-the-fly reconstruction practically feasible, enabling online 4D reconstruction. However,

Cited by 0SourceScholar
2026

MonoCLUE: Object-Aware Clustering Enhances Monocular 3D Object Detection

AAAI 2026technical

Monocular 3D object detection offers a cost-effective solution for autonomous driving, but it suffers from the ill-posed depth and a limited field of view. These constraints lead to the lack of geometric cues and reduced accuracy in occluded or truncated scenes. While recent approaches incorporate a

Cited by 0SourcePDFScholar
2025

CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images

CVPR 2025poster

3D Gaussian Splatting (3DGS) has attracted significant attention for its high-quality novel view rendering, inspiring research to address real-world challenges. While conventional methods depend on sharp images for accurate scene reconstruction, real-world scenarios are often affected by defocus blu…

Cited by 0SourcePDFScholar
2025

CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images

ICCV 2025poster

3D Gaussian Splatting (3DGS) has gained significant attention for their high-quality novel view rendering, motivating research to address real-world challenges. A critical issue is the camera motion blur caused by movement during exposure, which hinders accurate 3D scene reconstruction. In this stud…

2025

Effective SAM Combination for Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such as the Segment Anything Model (SAM), with a pre-trained vision-language model l…

Cited by 0SourcePDFScholar
2025

Elevating Flow-Guided Video Inpainting with Reference Generation

AAAI 2025technical

Video inpainting (VI) is a challenging task that requires effective propagation of observable content across frames while simultaneously generating new content not present in the original video. In this study, we propose a robust and practical VI framework that leverages a large generative model for…

2025

Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding

NeurIPS 2025poster

Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: \textit{Moment Retrieval (MR)} and \textit{Highlight Detection (HD)}. While recent advances have been progressed by…

Cited by 0SourceScholar
2025

Video Diffusion Models Are Strong Video Inpainter

AAAI 2025technical

Propagation-based video inpainting using optical flow at the pixel or feature level has recently garnered significant attention. However, it has limitations such as the inaccuracy of optical flow prediction and the propagation of noise over time. These issues result in non-uniform noise and time con…

Cited by 5SourcePDFScholar
2024

Dual Prototype Attention for Unsupervised Video Object Segmentation

CVPR 2024poster

Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel pro…

2024

FIMP: Future Interaction Modeling for Multi-Agent Motion Prediction

ICRA 2024poster

Multi-agent motion prediction is a crucial concern in autonomous driving, yet it remains a challenge owing to the ambiguous intentions of dynamic agents and their intricate interactions. Existing studies have attempted to capture interactions between road entities by using the definite data in histo…

Cited by 4SourceScholar
2024

Guided Slot Attention for Unsupervised Video Object Segmentation

CVPR 2024poster

Unsupervised video object segmentation aims to segment the most prominent object in a video sequence. However the existence of complex backgrounds and multiple foreground objects make this task challenging. To address this issue we propose a guided slot attention network to reinforce spatial structu…

2024

Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech Separation

NeurIPS 2024poster

In speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage of the network. Instead, we propose a more intuitive strateg…

2024

Towards Multi-Domain Learning for Generalizable Video Anomaly Detection

NeurIPS 2024poster

Most of the existing Video Anomaly Detection (VAD) studies have been conducted within single-domain learning, where training and evaluation are performed on a single dataset. However, the criteria for abnormal events differ across VAD datasets, making it problematic to apply a single-domain model to…

Cited by 1SourcePDFScholar
2023

DP-NeRF: Deblurred Neural Radiance Field With Physical Scene Priors

CVPR 2023poster

Neural Radiance Field (NeRF) has exhibited outstanding three-dimensional (3D) reconstruction quality via the novel view synthesis from multi-view images and paired calibrated camera parameters. However, previous NeRF-based systems have been demonstrated under strictly controlled settings, with littl…

2023

Exploring Discontinuity for Video Frame Interpolation

CVPR 2023highlight

Video frame interpolation (VFI) is the task that synthesizes the intermediate frame given two consecutive frames. Most of the previous studies have focused on appropriate frame warping operations and refinement modules for the warped frames. These studies have been conducted on natural videos contai…

2023

Exploring Temporally Dynamic Data Augmentation for Video Recognition

ICLR 2023top-25%

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been rarely explored despite its effectiveness. Few existing augmentation recipes for video recognition naively extend the im…

Cited by 13SourcePDFScholar
2023

FAPM: Fast Adaptive Patch Memory for Real-Time Industrial Anomaly Detection

ICASSP 2023accepted

Feature embedding-based methods have shown exceptional performance in detecting industrial anomalies by comparing features of target images with normal images. However, some methods do not meet the speed requirements of real-time inference, which is crucial for real-world applications. To address th…

Cited by 0SourceScholar
2023

Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action Recognition

ICCV 2023poster

Graph convolutional networks (GCNs) are the most commonly used methods for skeleton-based action recognition and have achieved remarkable performance. Generating adjacency matrices with semantically meaningful edges is particularly important for this task, but extracting such edges is challenging pr…

Cited by 222PDFcodeScholar
2023

Leveraging Spatio-Temporal Dependency for Skeleton-Based Action Recognition

ICCV 2023poster

Skeleton-based action recognition has attracted considerable attention due to its compact representation of the human body's skeletal sructure. Many recent methods have achieved remarkable performance using graph convolutional networks (GCNs) and convolutional neural networks (CNNs), which extract s…

Cited by 29PDFcodeScholar
2023

Look Around for Anomalies: Weakly-Supervised Anomaly Detection via Context-Motion Relational Learning

CVPR 2023poster

Weakly-supervised Video Anomaly Detection is the task of detecting frame-level anomalies using video-level labeled training data. It is difficult to explore class representative features using minimal supervision of weak labels with a single backbone branch. Furthermore, in real-world scenarios, the…

Cited by 49SourcePDFScholar
2023

Two-Stream Decoder Feature Normality Estimating Network for Industrial Anomaly Detection

ICASSP 2023accepted

Image reconstruction-based anomaly detection has recently been in the spotlight because of the difficulty of constructing anomaly datasets. These approaches work by learning to model normal features without seeing abnormal samples during training and then discriminating anomalies at test time based…

Cited by 0SourceScholar
2022

Expanded Adaptive Scaling Normalization for End to End Image Compression

ECCV 2022poster

"Recently, learning-based image compression methods that utilize convolutional neural layers have been developed rapidly. Rescaling modules such as batch normalization which are often used in convolutional neural networks do not operate adaptively for the various inputs. Therefore, Generalized Divis…

2022

Occluded Person Re-Identification Via Relational Adaptive Feature Correction Learning

ICASSP 2022accepted

Occluded person re-identification (Re-ID) in images captured by multiple cameras is challenging because the target person is occluded by pedestrians or objects, especially in crowded scenes. In addition to the processes performed during holistic person Re-ID, occluded person Re-ID involves the remov…

Cited by 0SourceScholar
2022

SPSN: Superpixel Prototype Sampling Network for RGB-D Salient Object Detection

ECCV 2022poster

"RGB-D salient object detection (SOD) has been in the spotlight recently because it is an important preprocessing operation for various vision tasks. However, despite advances in deep learning-based methods, RGB-D SOD is still challenging due to the large domain gap between an RGB image and the dept…

2022

Tackling Background Distraction in Video Object Segmentation

ECCV 2022poster

"Semi-supervised video object segmentation (VOS) aims to densely track certain designated objects in videos. One of the main challenges in this task is the existence of background distractors that appear similar to the target objects. We propose three novel strategies to suppress such distractors: 1…

2021

Regularization Strategy for Point Cloud via Rigidly Mixed Sample

CVPR 2021poster

Data augmentation is an effective regularization strategy to alleviate the overfitting, which is an inherent drawback of the deep neural networks. However, data augmentation is rarely considered for point cloud processing despite many studies proposing various augmentation methods for image data. Ac…

Cited by 101PDFcodeScholar
2020

AdaCoF: Adaptive Collaboration of Flows for Video Frame Interpolation

CVPR 2020poster

Video frame interpolation is one of the most challenging tasks in video processing research. Recently, many studies based on deep learning have been suggested. Most of these methods focus on finding locations with useful information to estimate each output pixel using their own frame warping operati…

Cited by 295PDFcodeScholar