← Search

Seon Joo Kim

54 accepted papers

2026

4D Scaffold Gaussian Splatting with Dynamic-Aware Anchor Growing for Efficient and High-Fidelity Dynamic Scene Reconstruction

AAAI 2026technical

Modeling dynamic scenes through 4D Gaussians offers high visual fidelity and fast rendering speeds, but comes with significant storage overhead. Recent approaches mitigate this cost by aggressively reducing the number of Gaussians. However, this inevitably removes Gaussians essential for high-qualit

Cited by 10SourcePDFScholar
2026

Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning Segmentation

ICLR 2026poster

Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to instructions. To exploit this for training-free localization, we cast video reasoning segmentation as video QA and extract attention maps via rollout. Since raw maps are too nois…

Cited by 0SourcecodeScholar
2026

Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

CVPR 2026

Most existing 3D keypoint estimation methods rely on manual annotations or calibrated multi-view images, both of which are expensive to collect.This paper introduces KeyDiff3D, a framework that can accurately predict 3D keypoints from a single image, thus eliminating the need for such expensive data

Cited by 0SourceScholar
2025

Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks

ICML 2025poster

Scaling has been a major driver of recent advancements in deep learning. Numerous empirical studies have found that scaling laws often follow the power-law and proposed several variants of power-law functions to predict the scaling behavior at larger scales. However, existing methods mostly rely on…

Cited by 0SourcePDFScholar
2025

CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text Matching

ICCV 2025poster

Text-to-video retrieval serves as a powerful tool for navigating vast video databases. This is particularly useful in autonomous driving to retrieve scenes from a text query to simulate and evaluate the driving system in desired scenarios. However, traditional ranking-based retrieval methods often r…

Cited by 0SourcePDFScholar
2025

CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color Constancy

ICCV 2025poster

Computational color constancy, or white balancing, is a key module in a camera's image signal processor (ISP) that corrects color casts from scene lighting. Because this operation occurs in the camera-specific raw color space, white balance algorithms must adapt to different cameras. This paper intr…

Cited by 0SourcePDFScholar
2025

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors

ICCV 2025poster

Recent advances in novel view synthesis (NVS) have enabled real-time rendering with 3D Gaussian Splatting (3DGS). However, existing methods struggle with artifacts and missing regions when rendering unseen viewpoints, limiting seamless scene exploration. To address this, we propose a 3DGS-based pipe…

Cited by 0SourcePDFScholar
2025

Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models

CVPR 2025poster

In this paper, we propose LSRNA, a novel framework for higher-resolution (exceeding 1K) image generation using diffusion models by leveraging super-resolution directly in the latent space. Existing diffusion models struggle with scaling beyond their training resolutions, often leading to structural…

2025

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

ICCV 2025poster

Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insigh…

2025

ORIDa: Object-centric Real-world Image Composition Dataset

CVPR 2025poster

Object compositing, the task of placing and harmonizing objects in images of diverse visual scenes, has become an important task in computer vision with the rise of generative models.However, existing datasets lack the diversity and scale required to comprehensively explore real-world scenarios comp…

Cited by 0SourcePDFScholar
2025

Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

CVPR 2025poster

We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we introduce Token Mark, a set of tokens highlighting the target regions within the vis…

Cited by 2SourcePDFScholar
2025

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

ICCV 2025poster

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained temporal annotations, we demonstrate that LLMs can effectively group atomic actio…

Cited by 0SourcePDFScholar
2025

Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion Models

ICCV 2025poster

Constructing a compressed latent space through a variational autoencoder (VAE) is the key for efficient 3D diffusion models. This paper introduces COD-VAE that encodes 3D shapes into a COmpact set of 1D latent vectors without sacrificing quality. COD-VAE introduces a two-stage autoencoder scheme to…

2025

Robust and Consistent Online Video Instance Segmentation via Instance Mask Propagation

AAAI 2025technical

Recent advancements in online Video Instance Segmentation (VIS) methods show notable performance improvements across benchmarks. However, the leading methods in the tracking-by-detection paradigm often result in temporally inconsistent predictions at both instance-level and pixel-level that lead to…

Cited by 0SourcePDFScholar
2025

Seam360GS: Seamless 360deg Gaussian Splatting from Real-World Omnidirectional Images

ICCV 2025poster

360deg visual content is widely shared on platforms such as YouTube and plays a central role in virtual reality, robotics, and autonomous navigation. However, consumer-grade dual-fisheye systems consistently yield imperfect panoramas due to inherent lens separation and angular distortions. In this w…

Cited by 0SourcePDFScholar
2025

UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations

CoRL 2025poster

Mimicry is a fundamental learning mechanism in humans, enabling individuals to learn new tasks by observing and imitating experts. However, applying this ability to robots presents significant challenges due to the inherent differences between human and robot embodiments in both their visual appeara…

Cited by 0SourceScholar
2024

Attentive Illumination Decomposition Model for Multi-Illuminant White Balancing

CVPR 2024poster

White balance (WB) algorithms in many commercial cameras assume single and uniform illumination leading to undesirable results when multiple lighting sources with different chromaticities exist in the scene. Prior research on multi-illuminant WB typically predicts illumination at the pixel level wit…

Cited by 6SourcePDFScholar
2023

A Generalized Framework for Video Instance Segmentation

CVPR 2023poster

The handling of long videos with complex and occluded sequences has recently emerged as a new challenge in the video instance segmentation (VIS) community. However, existing methods have limitations in addressing this challenge. We argue that the biggest bottleneck in current approaches is the discr…

2023

MiniROAD: Minimal RNN Framework for Online Action Detection

ICCV 2023poster

Online Action Detection (OAD) is the task of identifying actions in streaming videos without access to future frames. Much effort has been devoted to effectively capturing long-range dependencies, with transformers receiving the spotlight for their ability to capture long-range temporal structures.…

Cited by 22PDFcodeScholar
2023

Shepherding Slots to Objects: Towards Stable and Robust Object-Centric Learning

CVPR 2023poster

Object-centric learning (OCL) aspires general and com- positional understanding of scenes by representing a scene as a collection of object-centric representations. OCL has also been extended to multi-view image and video datasets to apply various data-driven inductive biases by utilizing geometric…

2023

Soft-Landing Strategy for Alleviating the Task Discrepancy Problem in Temporal Action Localization Tasks

CVPR 2023poster

Temporal Action Localization (TAL) methods typically operate on top of feature sequences from a frozen snippet encoder that is pretrained with the Trimmed Action Classification (TAC) tasks, resulting in a task discrepancy problem. While existing TAL methods mitigate this issue either by retraining t…

2022

Cannot See the Forest for the Trees: Aggregating Multiple Viewpoints To Better Classify Objects in Videos

CVPR 2022poster

Recently, both long-tailed recognition and object tracking have made great advances individually. TAO benchmark presented a mixture of the two, long-tailed object tracking, in order to further reflect the aspect of the real-world. To date, existing solutions have adopted detectors showing robustness…

Cited by 5PDFcodeScholar
2022

ComMU: Dataset for Combinatorial Music Generation

NeurIPS 2022accept

Commercial adoption of automatic music composition requires the capability of generating diverse and high-quality music suitable for the desired context (e.g., music for romantic movies, action games, restaurants, etc.). In this paper, we introduce combinatorial music generation, a new task to creat…

2022

UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection

CVPR 2022poster

Generic Event Boundary Detection (GEBD) is a newly suggested video understanding task that aims to find one level deeper semantic boundaries of events. Bridging the gap between natural human perception and video understanding, it has various potential applications, including interpretable and semant…

Cited by 36PDFcodeScholar
2022

VISOLO: Grid-Based Space-Time Aggregation for Efficient Online Video Instance Segmentation

CVPR 2022oral

For online video instance segmentation (VIS), fully utilizing the information from previous frames in an efficient manner is essential for real-time applications. Most previous methods follow a two-stage approach requiring additional computations such as RPN and RoIAlign, and do not fully exploit th…

Cited by 42PDFcodeScholar
2022

VITA: Video Instance Segmentation via Object Token Association

NeurIPS 2022accept

We introduce a novel paradigm for offline Video Instance Segmentation (VIS), based on the hypothesis that explicit object-oriented information can be a strong clue for understanding the context of the entire sequence. To this end, we propose VITA, a simple structure built on top of an off-the-shelf…

2021

CAG-QIL: Context-Aware Actionness Grouping via Q Imitation Learning for Online Temporal Action Localization

ICCV 2021poster

Temporal action localization has been one of the most popular tasks in video understanding, due to the importance of detecting action instances in videos. However, not much progress has been made on extending it to work in an online fashion, although many video related tasks can benefit by going onl…

Cited by 14PDFScholar
2021

Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm Under Mixed Illumination

ICCV 2021poster

We introduce a Large Scale Multi-Illuminant (LSMI) Dataset that contains 7,486 images, captured with three different cameras on more than 2,700 scenes with two or three illuminants. For each image in the dataset, the new dataset provides not only the pixel-wise ground truth illumination but also the…

Cited by 31PDFcodeScholar
2021

Tackling the Ill-Posedness of Super-Resolution Through Adaptive Target Generation

CVPR 2021poster

By the one-to-many nature of the super-resolution (SR) problem, a single low-resolution (LR) image can be mapped to many high-resolution (HR) images. However, learning based SR algorithms are trained to map an LR image to the corresponding ground truth (GT) HR image in the training dataset. The trai…

Cited by 62PDFcodeScholar
2021

Video Instance Segmentation using Inter-Frame Communication Transformers

NeurIPS 2021poster

We propose a novel end-to-end solution for video instance segmentation (VIS) based on transformers. Recently, the per-clip pipeline shows superior performance over per-frame methods leveraging richer information from multiple frames. However, previous per-clip models require heavy computation and…

2020

Cross-Identity Motion Transfer for Arbitrary Objects through Pose-Attentive Video Reassembling

ECCV 2020poster

We propose an attention-based networks for transferring motions between arbitrary objects. Given a source image(s) and a driving video, our networks animate the subject in the source images according to the motion in the driving video. In our attention mechanism, dense similarities between the learn…

Cited by 13SourcePDFScholar
2020

Deep Space-Time Video Upsampling Networks

ECCV 2020poster

Video super-resolution (VSR) and frame interpolation (FI) are traditional computer vision problems, and the performance have been improving by incorporating deep learning recently. In this paper, we investigate the problem of jointly upsampling videos both in space and time, which is becoming more i…

2019

End-To-End Time-Lapse Video Synthesis From a Single Outdoor Image

CVPR 2019poster

Time-lapse videos usually contain visually appealing content but are often difficult and costly to create. In this paper, we present an end-to-end solution to synthesize a time-lapse video from a single outdoor image using deep neural networks. Our key idea is to train a conditional generative adver…

Cited by 40PDFScholar
2019

Fast User-Guided Video Object Segmentation by Interaction-And-Propagation Networks

CVPR 2019poster

We present a deep learning method for the interactive video object segmentation. Our method is built upon two core operations, interaction and propagation, and each operation is conducted by Convolutional Neural Networks. The two networks are connected both internally and externally so that the netw…

Cited by 79PDFScholar
2019

Unsupervised Keypoint Learning for Guiding Class-Conditional Video Prediction

NeurIPS 2019poster

We propose a deep video prediction model conditioned on a single image and an action class. To generate future frames, we first detect keypoints of a moving object and predict future motion as a sequence of keypoints. The input image is then translated following the predicted keypoints sequence to c…

Cited by 58SourcePDFScholar
2018

Deep Video Super-Resolution Network Using Dynamic Upsampling Filters Without Explicit Motion Compensation

CVPR 2018poster

Video super-resolution (VSR) has become even more important recently to provide high resolution (HR) contents for ultra high definition displays. While many deep learning based VSR methods have been proposed, most of them rely heavily on the accuracy of motion estimation and compensation. We introdu…

2018

EPINET: A Fully-Convolutional Neural Network Using Epipolar Geometry for Depth From Light Field Images

CVPR 2018poster

Light field cameras capture both the spatial and the angular properties of light rays in space. Due to its property, one can compute the depth from light fields in uncontrolled lighting environments, which is a big advantage over active sensing devices. Depth computed from light fields can be used f…

Cited by 325SourcePDFScholar
2018

Fast Video Object Segmentation by Reference-Guided Mask Propagation

CVPR 2018poster

We present an efficient method for the semi-supervised video object segmentation. Our method achieves accuracy competitive with state-of-the-art methods while running in a fraction of time compared to others. To this end, we propose a deep Siamese encoder-decoder network that is designed to take adv…

Cited by 512SourcePDFScholar
2018

Teaching Machines to Understand Baseball Games: Large-Scale Baseball Video Database for Multiple Video Understanding Tasks

ECCV 2018poster

A major obstacle in teaching machines to understand videos is the lack of training data, as creating temporal annotations for long videos requires a huge amount of human effort. To this end, we introduce a new large-scale baseball video dataset called the BBDB, which is produced semi-automatically b…

Cited by 16SourcePDFScholar
2018

Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language

NeurIPS 2018spotlight

This paper addresses the problem of manipulating images using natural language description. Our task aims to semantically modify visual attributes of an object in an image according to the text describing the new visual appearance. Although existing methods synthesize images having new attributes, t…

Cited by 263SourcePDFScholar
2016

A Holistic Approach to Cross-Channel Image Noise Modeling and Its Application to Image Denoising

CVPR 2016spotlight

Modelling and analyzing noise in images is a fundamental task in many computer vision systems. Traditionally, noise has been modelled per color channel assuming that the color channels are independent. Although the color channels can be considered as mutually independent in camera RAW images, signal…

Cited by 310PDFScholar
2016

Do It Yourself Hyperspectral Imaging With Everyday Digital Cameras

CVPR 2016spotlight

Capturing hyperspectral images requires expensive and specialized hardware that is not readily accessible to most users. Digital cameras, on the other hand, are significantly cheaper in comparison and can be easily purchased and used. In this paper, we present a framework for reconstructing hyperspe…

Cited by 110PDFScholar
2015

Complementary Sets of Shutter Sequences for Motion Deblurring

ICCV 2015poster

In this paper, we present a novel multi-image motion deblurring method utilizing the coded exposure technique. The key idea of our work is to capture video frames with a set of complementary fluttering patterns to preserve spatial frequency details. We introduce an algorithm for generating a complem…

Cited by 7PDFScholar