← Search

Yujin Jeong

5 accepted papers

2026

When Do Diffusion Models learn to Generate Multiple Objects?

ICML 2026poster

Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation. Despite extensive empirical evidence of these failures, the underlying causes remain unclear. We begin by asking how much of this limitation arises from the data itself. To disen…

Cited by 0SourceScholar
2025

Diffusion Classifiers Understand Compositionality, but Conditions Apply

NeurIPS 2025poster

Understanding visual scenes is fundamental to human intelligence. While discriminative models have significantly advanced computer vision, they often struggle with compositional understanding. In contrast, recent generative text-to-image diffusion models excel at synthesizing complex scenes, suggest…

Cited by 0SourcecodeScholar
2025

Read, Watch and Scream! Sound Generation from Text and Video

AAAI 2025technical

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely, text-to-audio generation methods generate high-quality audio b…

2025

Zero-Shot Compositional Video Learning with Coding Rate Reduction

ICCV 2025poster

In this paper, we propose a novel zero-shot compositional video understanding method inspired by how young children efficiently learn new concepts and flexibly expand their existing knowledge framework. While recent large-scale visual language models (VLMs) have achieved remarkable advancements and…

2023

The Power of Sound (TPoS): Audio Reactive Video Generation with Stable Diffusion

ICCV 2023poster

In recent years, video generation has become a prominent generative tool and has drawn significant attention. However, there is little consideration in audio-to-video generation, though audio contains unique qualities like temporal semantics and magnitude. Hence, we propose The Power of Sound (TPoS)…

Cited by 38PDFcodeScholar