← Search

Youngjung Uh

24 accepted papers

2026

4D Scaffold Gaussian Splatting with Dynamic-Aware Anchor Growing for Efficient and High-Fidelity Dynamic Scene Reconstruction

AAAI 2026technical

Modeling dynamic scenes through 4D Gaussians offers high visual fidelity and fast rendering speeds, but comes with significant storage overhead. Recent approaches mitigate this cost by aggressively reducing the number of Gaussians. However, this inevitably removes Gaussians essential for high-qualit

Cited by 10SourcePDFScholar
2026

MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion

ICLR 2026poster

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with geometric consistency, whereas customization models lack explicit…

Cited by 0SourceScholar
2026

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

ICLR 2026poster

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a promising condition for temporally controlled video generatio…

Cited by 0SourceScholar
2025

Rethinking Open-Vocabulary Segmentation of Radiance Fields in 3D Space

AAAI 2025technical

Understanding the 3D semantics of a scene is a fundamental problem for various scenarios such as embodied agents. While NeRFs and 3DGS excel at novel-view synthesis, previous methods for understanding their semantics have been limited to incomplete 3D understanding: their segmentation results are re…

Cited by 2SourcePDFScholar
2025

StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance

ICCV 2025poster

In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled more precise control over style and content. However, existing methods often suffer from content leakage, where undesired…

Cited by 0SourcePDFScholar
2025

TCFG: Tangential Damping Classifier-free Guidance

CVPR 2025poster

Diffusion models have achieved remarkable success in text-to-image synthesis, largely attributed to the use of classifier-free guidance (CFG), which enables high-quality, condition-aligned image generation. CFG combines the conditional score (e.g., text-conditioned) with the unconditional score to c…

Cited by 0SourcePDFScholar
2024

Attribute Based Interpretable Evaluation Metrics for Generative Models

ICML 2024poster

When the training dataset comprises a 1:1 proportion of dogs to cats, a generative model that produces 1:1 dogs and cats better resembles the training species distribution than another model with 3:1 dogs and cats. Can we capture this phenomenon using existing metrics? Unfortunately, we cannot, beca…

2024

Sync-NeRF: Generalizing Dynamic NeRFs to Unsynchronized Videos

AAAI 2024technical

Recent advancements in 4D scene reconstruction using neural radiance fields (NeRF) have demonstrated the ability to represent dynamic scenes from multi-view videos. However, they fail to reconstruct the dynamic scenes and struggle to fit even the training views in unsynchronized settings. It happens…

2023

AesPA-Net: Aesthetic Pattern-Aware Style Transfer Networks

ICCV 2023poster

To deliver the artistic expression of the target style, recent studies exploit the attention mechanism owing to its ability to map the local patches of the style image to the corresponding patches of the content image. However, because of the low semantic correspondence between arbitrary content and…

Cited by 42PDFcodeScholar
2023

BallGAN: 3D-aware Image Synthesis with a Spherical Background

ICCV 2023poster

3D-aware GANs aim to synthesize realistic 3D scenes that can be rendered in arbitrary camera viewpoints, generating high-quality images with well-defined geometry. As 3D content creation becomes more popular, the ability to generate foreground objects separately from the background has become a cruc…

Cited by 7PDFScholar
2023

LANIT: Language-Driven Image-to-Image Translation for Unlabeled Data

CVPR 2023poster

Existing techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent truly-unsupervised methods adopt clustering approaches to easily provide per-sample on…

2023

Semantic Image Synthesis with Unconditional Generator

NeurIPS 2023poster

Semantic image synthesis (SIS) aims to generate realistic images according to semantic masks given by a user. Although recent methods produce high quality results with fine spatial control, SIS requires expensive pixel-level annotation of the training images. On the other hand, manipulating intermed…

Cited by 4SourcePDFScholar
2023

Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry

NeurIPS 2023poster

Despite the success of diffusion models (DMs), we still lack a thorough understanding of their latent space. To understand the latent space $\mathbf{x}_t \in \mathcal{X}$, we analyze them from a geometrical perspective. Our approach involves deriving the local latent basis within $\mathcal{X}$ by le…

2021

AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights

ICLR 2021poster

Normalization techniques, such as batch normalization (BN), are a boon for modern deep learning. They let weights converge more quickly with often better generalization performances. It has been argued that the normalization-induced scale invariance among the weights provides an advantageous ground…

2021

Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image Editing

CVPR 2021poster

Generative adversarial networks (GANs) synthesize realistic images from random latent vectors. Although manipulating the latent vectors controls the synthesized outputs, editing real images with GANs suffers from i) time-consuming optimization for projecting real images to the latent vectors, ii) or…

Cited by 192PDFcodeScholar
2021

Rethinking the Truly Unsupervised Image-to-Image Translation

ICCV 2021poster

Every recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image transl…

Cited by 120PDFcodeScholar
2020

Reliable Fidelity and Diversity Metrics for Generative Models

ICML 2020poster

Devising indicative evaluation metrics for the image generation task remains an open problem. The most widely used metric for measuring the similarity between real and generated images has been the Frechet Inception Distance (FID) score. Since it does not differentiate the fidelity and diversity asp…

2019

Photorealistic Style Transfer via Wavelet Transforms

ICCV 2019poster

Recent style transfer models have provided promising artistic results. However, given a photograph as a reference style, existing methods are limited by spatial distortions or unrealistic artifacts, which should not happen in real photographs. We introduce a theoretically sound correction to the net…

Cited by 425PDFcodeScholar