← Search

Joon-Young Lee

61 accepted papers

2026

AnthroTAP: Learning Point Tracking with Real-World Motion

CVPR 2026

Point tracking models often struggle to generalize to real-world videos because large-scale training data is predominantly synthetic--the only source currently feasible to produce at scale. Collecting real-world annotations, however, is prohibitively expensive, as it requires tracking hundreds of po

Cited by 0SourcecodeScholar
2026

DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation

CVPR 2026

Estimating accurate, view-consistent geometry and camera poses from uncalibrated multi-view/video inputs remains challenging--especially at high spatial resolutions and over long sequences. We present DAGE, a dual-stream transformer whose main novelty is to disentangle global coherence from fine det

Cited by 0SourcecodeScholar
2026

Generative Video Motion Editing with 3D Point Tracks

CVPR 2026

Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video edi

Cited by 0SourceScholar
2026

VideoMaMa: Mask-Guided Video Matting via Generative Prior

CVPR 2026

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model VideoMaMa that converts coarse segmentation masks into pixel accurate alpha mattes, by leveraging pretrained video diffusi

Cited by 0SourcecodeScholar
2025

Elevating Flow-Guided Video Inpainting with Reference Generation

AAAI 2025technical

Video inpainting (VI) is a challenging task that requires effective propagation of observable content across frames while simultaneously generating new content not present in the original video. In this study, we propose a robust and practical VI framework that leverages a large generative model for…

2025

Exploring Temporally-Aware Features for Point Tracking

CVPR 2025poster

Point tracking in videos is a fundamental task with applications in robotics, video editing, and more. While many vision tasks benefit from pre-trained feature backbones to improve generalizability, point tracking has primarily relied on simpler backbones trained from scratch on synthetic data, whic…

2025

Generative Video Propagation

CVPR 2025poster

Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of su…

Cited by 1SourcePDFScholar
2025

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

ICCV 2025poster

Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with token count. To address this, we propose a training-free spatio-temporal token merging method, named STTM. Our key insigh…

2025

Robust and Consistent Online Video Instance Segmentation via Instance Mask Propagation

AAAI 2025technical

Recent advancements in online Video Instance Segmentation (VIS) methods show notable performance improvements across benchmarks. However, the leading methods in the tracking-by-detection paradigm often result in temporally inconsistent predictions at both instance-level and pixel-level that lead to…

Cited by 0SourcePDFScholar
2025

Video Color Grading via Look-Up Table Generation

ICCV 2025poster

Different from color correction and transfer, color grading involves adjusting colors for artistic or storytelling purposes in a video, which is used to establish a specific look or mood. However, due to the complexity of the process and the need for specialized editing skills, video color grading r…

2024

Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

CVPR 2024poster

While there has been significant progress in customizing text-to-image generation models generating images that combine multiple personalized concepts remains challenging. In this work we introduce Concept Weaver a method for composing customized text-to-image diffusion models at inference time. Spe…

Cited by 12SourcePDFScholar
2024

FlowTrack: Revisiting Optical Flow for Long-Range Dense Tracking

CVPR 2024poster

In the domain of video tracking existing methods often grapple with a trade-off between spatial density and temporal range. Current approaches in dense optical flow estimators excel in providing spatially dense tracking but are limited to short temporal spans. Conversely recent advancements in long-…

Cited by 9SourcePDFScholar
2024

HARIVO: Harnessing Text-to-Image Models for Video Generation

ECCV 2024poster

"We present a method to create diffusion-based video models from pretrained Text-to-Image (T2I) models. Recently, AnimateDiff proposed freezing the T2I model while only training temporal layers. We advance this method by proposing a unique architecture, incorporating a mapping network and frame-wise…

2024

MaGGIe: Masked Guided Gradual Human Instance Matting

CVPR 2024poster

Human matting is a foundation task in image and video processing where human foreground pixels are extracted from the input. Prior works either improve the accuracy by additional guidance or improve the temporal consistency of a single instance across frames. We propose a new framework MaGGIe Masked…

2024

Putting the Object Back into Video Object Segmentation

CVPR 2024highlight

We present Cutie a video object segmentation (VOS) network with object-level memory reading which puts the object representation from memory back into the video object segmentation result. Recent works on VOS employ bottom-up pixel-level memory reading which struggles due to matching noise especiall…

2023

A Generalized Framework for Video Instance Segmentation

CVPR 2023poster

The handling of long videos with complex and occluded sequences has recently emerged as a new challenge in the video instance segmentation (VIS) community. However, existing methods have limitations in addressing this challenge. We argue that the biggest bottleneck in current approaches is the discr…

2023

Long-range Multimodal Pretraining for Movie Understanding

ICCV 2023poster

Learning computer vision models from (and for) movies has a long-standing history. While great progress has been attained, there is still a need for a pretrained multimodal model that can perform well in the ever-growing set of movie understanding tasks the community has been establishing. In this w…

Cited by 16PDFScholar
2023

Tracking Anything with Decoupled Video Segmentation

ICCV 2023poster

Training data for video segmentation are expensive to annotate. This impedes extensions of end-to-end algorithms to new video segmentation tasks, especially in large-vocabulary settings. To 'track anything' without training on video data for every individual task, we develop a decoupled video segmen…

Cited by 269PDFcodeScholar
2023

XMem++: Production-level Video Segmentation From Few Annotated Frames

ICCV 2023poster

Despite advancements in user-guided video segmentation, extracting complex objects consistently for highly complex scenes is still a labor-intensive task, especially for production. It is not uncommon that a majority of frames need to be annotated. We introduce a novel semi-supervised video object s…

Cited by 33PDFcodeScholar
2022

Bridging Images and Videos: A Simple Learning Framework for Large Vocabulary Video Object Detection

ECCV 2022poster

"Scaling object taxonomies is one of the important steps toward a robust real-world deployment of recognition systems. We have faced remarkable progress in images since the introduction of the LVIS benchmark. To continue this success in videos, a new video benchmark, TAO, was recently presented. Giv…

Cited by 8SourcePDFScholar
2022

Information-Theoretic Bias Reduction via Causal View of Spurious Correlation

AAAI 2022technical

We propose an information-theoretic bias measurement technique through a causal interpretation of spurious correlation, which is effective to identify the feature-level algorithmic bias by taking advantage of conditional mutual information. Although several bias measurement methods have been propose…

Cited by 26SourcePDFScholar
2022

The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video Editing

ECCV 2022poster

"Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, color grading, or applying digital makeups. However, most of the solutions have focused on video manipulation and VFX. This…

2022

VITA: Video Instance Segmentation via Object Token Association

NeurIPS 2022accept

We introduce a novel paradigm for offline Video Instance Segmentation (VIS), based on the hypothesis that explicit object-oriented information can be a strong clue for understanding the context of the entire sequence. To this end, we propose VITA, a simple structure built on top of an off-the-shelf…

2021

Hierarchical Memory Matching Network for Video Object Segmentation

ICCV 2021poster

We present Hierarchical Memory Matching Network (HMMN) for semi-supervised video object segmentation. Based on a recent memory-based method [33], we propose two advanced memory read modules that enable us to perform memory reading in multiple scales while exploiting temporal smoothness. We first pro…

Cited by 149PDFcodeScholar
2020

Active Speakers in Context

CVPR 2020poster

Current methods for active speaker detection focus on modeling audiovisual information from a single speaker. This strategy can be adequate for addressing single-speaker scenarios, but it prevents accurate detection when the task is to identify who of many candidate speakers are talking. This paper…

Cited by 104PDFcodeScholar
2020

Learning Visual Emotion Representations From Web Data

CVPR 2020poster

We present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmoti…

Cited by 50PDFScholar
2020

URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

ECCV 2020poster

We propose a unified referring video object segmentation network (URVOS). URVOS takes a video and a referring expression as inputs, and estimates the {object masks} referred by the given language expression in the whole video frames. Our algorithm addresses the challenging problem by performing lang…

2019

Fast User-Guided Video Object Segmentation by Interaction-And-Propagation Networks

CVPR 2019poster

We present a deep learning method for the interactive video object segmentation. Our method is built upon two core operations, interaction and propagation, and each operation is conducted by Convolutional Neural Networks. The two networks are connected both internally and externally so that the netw…

Cited by 79PDFScholar
2019

GAPLE: Generalizable Approaching Policy LEarning for Robotic Object Searching in Indoor Environment

RA-L 2019

We study the problem of learning a generalizable action policy for an intelligent agent to actively approach an object of interest, in an indoor environment, solely from its visual inputs. While scene-driven or recognition-driven visual navigation has been widely studied, prior efforts suffer severe

Cited by 20SourceScholar
2018

Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias

ECCV 2018poster

While machine learning approaches to visual emotion recognition offer great promise, current methods consider training and testing models on small scale datasets covering limited visual emotion concepts. Our analysis identifies an important but long overlooked issue of existing visual emotion benchm…

Cited by 106SourcePDFScholar
2018

Distort-and-Recover: Color Enhancement Using Deep Reinforcement Learning

CVPR 2018poster

Learning-based color enhancement approaches typically learn to map from input images to retouched images. Most of existing methods require expensive pairs of input-retouched images or produce results in a non-interpretable way. In this paper, we present a deep reinforcement learning (DRL) based meth…

Cited by 261SourcePDFScholar
2018

Fast Video Object Segmentation by Reference-Guided Mask Propagation

CVPR 2018poster

We present an efficient method for the semi-supervised video object segmentation. Our method achieves accuracy competitive with state-of-the-art methods while running in a fraction of time compared to others. To this end, we propose a deep Siamese encoder-decoder network that is designed to take adv…

Cited by 512SourcePDFScholar
2018

Learning to Blend Photos

ECCV 2018poster

Photo blending is a common technique to create aesthetically pleasing artworks by combining multiple photos. However, the process of photo blending is usually time-consuming, and care must be taken in the process of blending, filtering, positioning, and masking each of the source photos. To make pho…

2018

RANUS: RGB and NIR Urban Scene Dataset for Deep Scene Parsing

RA-L 2018

In this letter, we present a data-driven method for scene parsing of road scenes to utilize single-channel near-infrared (NIR) images. To overcome the lack of data problem in non-RGB spectrum, we define a new color space and decompose the task of deep scene parsing into two subtasks with two separat

Cited by 42SourceScholar
2018

What do I Annotate Next? An Empirical Study of Active Learning for Action Localization

ECCV 2018poster

Despite tremendous progress achieved in temporal action localization, state-of-the-art methods still struggle to train accurate models when annotated data is scarce. In this paper, we introduce a novel active learning framework for temporal localization that aims to mitigate this data dependency iss…

Cited by 51SourcePDFScholar
2017

Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks

CVPR 2017poster

Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their representation learning capabilities. One of the bottlenecks in training…

Cited by 329PDFScholar
2017

Reflectance Capture Using Univariate Sampling of BRDFs

ICCV 2017poster

We propose the use of a light-weight setup consisting of a collocated camera and light source --- commonly found on mobile devices --- to reconstruct surface normals and spatially-varying BRDFs of near-planar material samples. A collocated setup provides only a 1-D "univariate" sampling of the 4-D B…

Cited by 73PDFScholar
2016

Automatic Content-Aware Color and Tone Stylization

CVPR 2016spotlight

We introduce a new technique that automatically generates diverse, visually compelling stylizations for a photograph in an unsupervised manner. We achieve this by learning style ranking for a given input using a large photo collection and selecting a diverse subset of matching styles for final style…

Cited by 92PDFScholar
2016

Stereo Matching With Color and Monochrome Cameras in Low-Light Conditions

CVPR 2016poster

Consumer devices with stereo cameras have become popular because of their low-cost depth sensing capability. However, those systems usually suffer from low imaging quality and inaccurate depth acquisition under low-light conditions. To address the problem, we present a new stereo matching method wit…

Cited by 60PDFScholar
2016

Vision system and depth processing for DRC-HUBO+

ICRA 2016

This paper presents a vision system and a depth processing algorithm for DRC-HUBO+, the winner of the DRC finals 2015. Our system is designed to reliably capture 3D information of a scene and objects and to be robust to challenging environment conditions. We also propose a depth-map upsampling metho

Cited by 13SourceScholar
2015

AttentionNet: Aggregating Weak Directions for Accurate Object Detection

ICCV 2015poster

We present a novel detection method using a deep convolutional neural network (CNN), named AttentionNet. We cast an object detection problem as an iterative classification problem, which is the most suitable form of a CNN. AttentionNet provides quantized weak directions pointing a target object and…

Cited by 236PDFcodeScholar
2015

Complementary Sets of Shutter Sequences for Motion Deblurring

ICCV 2015poster

In this paper, we present a novel multi-image motion deblurring method utilizing the coded exposure technique. The key idea of our work is to capture video frames with a set of complementary fluttering patterns to preserve spatial frequency details. We introduce an algorithm for generating a complem…

Cited by 7PDFScholar