← Search

Susung Hong

10 accepted papers

2026

MusicInfuser: Making Video Diffusion Listen and Dance

CVPR 2026

We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion model from scratch, our method demonstrates how existing video d

Cited by 0SourcecodeScholar
2026

TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling

ICML 2026poster

Recent diffusion models achieve the state-of-the-art performance in image generation, but often suffer from semantic inconsistencies or *hallucinations*. While various inference-time guidance methods can enhance generation, they often operate *indirectly* by relying on external signals or architectu…

Cited by 0SourceScholar
2025

Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories

CVPR 2025poster

Recent advancements in text-based diffusion models have accelerated progress in 3D reconstruction and text-based 3D editing. Although existing 3D editing methods excel at modifying color, texture, and style, they struggle with extensive geometric or appearance changes, thus limiting their applicatio…

2025

Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling

CVPR 2025poster

Diffusion models have emerged as a powerful tool for generating high-quality images, videos, and 3D content. While sampling guidance techniques like CFG improve quality, they reduce diversity and motion. Autoguidance mitigates these issues but demands extra weak model training, limiting its practica…

2024

Effective Rank Analysis and Regularization for Enhanced 3D Gaussian Splatting

NeurIPS 2024poster

3D reconstruction from multi-view images is one of the fundamental challenges in computer vision and graphics. Recently, 3D Gaussian Splatting (3DGS) has emerged as a promising technique capable of real-time rendering with high-quality 3D reconstruction. This method utilizes 3D Gaussian representati…

2024

Retrieval-Augmented Score Distillation for Text-to-3D Generation

ICML 2024poster

Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have been released, fine-tuning the diffusion model on the multi-v…

2024

Smoothed Energy Guidance: Guiding Diffusion Models with Reduced Energy Curvature of Attention

NeurIPS 2024poster

Conditional diffusion models have shown remarkable success in visual content generation, producing high-quality samples across various domains, largely due to classifier-free guidance (CFG). Recent attempts to extend guidance to unconditional models have relied on heuristic techniques, resulting in…

2023

Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D Generation

NeurIPS 2023poster

Existing score-distilling text-to-3D generation techniques, despite their considerable promise, often encounter the view inconsistency problem. One of the most notable issues is the Janus problem, where the most canonical view of an object (\textit{e.g}., face or head) appears in other views. In thi…

2023

Improving Sample Quality of Diffusion Models Using Self-Attention Guidance

ICCV 2023poster

Denoising diffusion models (DDMs) have attracted attention for their exceptional generation quality and diversity. This success is largely attributed to the use of class- or text-conditional diffusion guidance methods, such as classifier and classifier-free guidance. In this paper, we present a more…

Cited by 96PDFScholar
2022

Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence

NeurIPS 2022accept

Existing pipelines of semantic correspondence commonly include extracting high-level semantic features for the invariance against intra-class variations and background clutters. This architecture, however, inevitably results in a low-resolution matching field that additionally requires an ad-hoc int…