← Search

Sanghyun Woo

23 accepted papers

2026

VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement

CVPR 2026

Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. In this paper, we bridge this critical gap by tackling three key challenges in 3D object arrangement task using MLLMs. Fi

Cited by 0SourceScholar
2025

Epsilon-VAE: Denoising as Visual Decoding

ICML 2025poster

In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data, it reduces redundancy and emphasizes key features for high-quality generation. Current visual tokenization methods rely…

Cited by 0SourcePDFScholar
2024

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

NeurIPS 2024oral

We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision components are often insufficiently explored and disconnected from visual representation learning re…

2024

MTMMC: A Large-Scale Real-World Multi-Modal Camera Tracking Benchmark

CVPR 2024poster

Multi-target multi-camera tracking is a crucial task that involves identifying and tracking individuals over time using video streams from multiple cameras. This task has practical applications in various fields such as visual surveillance crowd behavior analysis and anomaly detection. However due t…

Cited by 1SourcePDFScholar
2024

SwitchLight: Co-design of Physics-driven Architecture and Pre-training Framework for Human Portrait Relighting

CVPR 2024highlight

We introduce a co-designed approach for human portrait relighting that combines a physics-guided architecture with a pre-training framework. Drawing on the Cook-Torrance reflectance model we have meticulously configured the architecture design to precisely simulate light-surface interactions. Furthe…

Cited by 19SourcePDFScholar
2023

Bidirectional Domain Mixup for Domain Adaptive Semantic Segmentation

AAAI 2023technical

Mixup provides interpolated training samples and allows the model to obtain smoother decision boundaries for better generalization. The idea can be naturally applied to the domain adaptation task, where we can mix the source and target samples to obtain domain-mixed samples for better adaptation. Ho…

2023

ConvNeXt V2: Co-Designing and Scaling ConvNets With Masked Autoencoders

CVPR 2023poster

Driven by improved architectures and better representation learning frameworks, the field of visual recognition has enjoyed rapid modernization and performance boost in the early 2020s. For example, modern ConvNets, represented by ConvNeXt models, have demonstrated strong performance across differen…

2023

Test-Time Adaptation in the Dynamic World With Compound Domain Knowledge Management

RA-L 2023

Prior to the deployment of robotic systems, pre-training the deep-recognition models on all potential visual cases is infeasible in practice. Hence, test-time adaptation (TTA) allows the model to adapt itself to novel environments and improve its performance during test time (i.e., lifelong adaptati

Cited by 9SourceScholar
2022

Bridging Images and Videos: A Simple Learning Framework for Large Vocabulary Video Object Detection

ECCV 2022poster

"Scaling object taxonomies is one of the important steps toward a robust real-world deployment of recognition systems. We have faced remarkable progress in images since the introduction of the LVIS benchmark. To continue this success in videos, a new video benchmark, TAO, was recently presented. Giv…

Cited by 8SourcePDFScholar
2021

LabOR: Labeling Only if Required for Domain Adaptive Semantic Segmentation

ICCV 2021poster

Unsupervised Domain Adaptation (UDA) for semantic segmentation has been actively studied to mitigate the domain gap between label-rich source data and unlabeled target data. Despite these efforts, UDA still has a long way to go to reach the fully supervised performance. To this end, we propose a Lab…

Cited by 55PDFScholar
2020

Discover, Hallucinate, and Adapt: Open Compound Domain Adaptation for Semantic Segmentation

NeurIPS 2020poster

Unsupervised domain adaptation (UDA) for semantic segmentation has been attracting attention recently, as it could be beneficial for various label-scarce real-world scenarios (e.g., robot control, autonomous driving, medical imaging, etc.). Despite the significant progress in this field, current wor…

Cited by 39SourcePDFScholar
2020

Global-and-Local Relative Position Embedding for Unsupervised Video Summarization

ECCV 2020poster

In order to summarize a content video properly, it is important to grasp the sequential structure of video as well as the long-term dependency between frames. The necessity of them is more obvious, especially for unsupervised learning. One possible solution is to utilize a well-known technique in th…

Cited by 75SourcePDFScholar
2020

Two-phase Pseudo Label Densification for Self-training based Domain Adaptation

ECCV 2020poster

Recently, deep self-training approaches emerged as a powerful solution to the unsupervised domain adaptation. The self-training scheme involves iterative processing of target data; it generates target pseudo labels and retrains the network. However, since only the confident predictions are taken as…

Cited by 132SourcePDFScholar