← Search

Sunjae Yoon

23 accepted papers

2026

Diffusion Negative Preference Optimization Made Simple

ICLR 2026poster

Classifier-Free Guidance (CFG) improves diffusion sampling by encouraging conditional generations while discouraging unconditional ones. Existing preference alignment methods, however, focus only on positive preference pairs, limiting their ability to actively suppress undesirable outputs. Diffusion…

Cited by 0SourcecodeScholar
2026

GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian Splatting

ICML 2026poster

Gaussian Splatting has achieved significant improvements by incorporating warping-based techniques. These approaches enhance synthesis quality by warping images from source views into the target viewpoint to compensate for missing or residual pixels. However, such methods suffer from pixel-level ina…

Cited by 0SourceScholar
2026

GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment Retrieval

AAAI 2026technical

Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and

Cited by 0SourcePDFScholar
2025

FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

ICML 2025spotlight

Drag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts…

Cited by 0SourcePDFScholar
2025

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

CVPR 2025poster

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion Transformer (MDT) for improved handling of both global garment cont…

Cited by 0SourcePDFScholar
2025

Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning

EMNLP 2025

Automated Audio Captioning (AAC) aims to generate natural language descriptions of audio content, enabling machines to interpret and communicate complex acoustic scenes. However, current AAC datasets often suffer from short and simplistic captions, limiting model expressiveness and semantic depth. T

Cited by 0SourcePDFScholar
2025

Occlusion-robust Stylization for Drawing-based 3D Animation

ICCV 2025poster

3D animation aims to generate a 3D animated video from an input image and a target 3D motion sequence. Recent advances in image-to-3D models enable the creation of animations directly from user-hand drawings. Distinguished from conventional 3D animation, drawing-based 3D animation is crucial to pres…

Cited by 0SourcePDFScholar
2024

DNI: Dilutional Noise Initialization for Diffusion Video Editing

ECCV 2024poster

"Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving the original structure of the input video. This limitation…

Cited by 2SourcePDFScholar
2024

FRAG: Frequency Adapting Group for Diffusion Video Editing

ICML 2024poster

In video editing, the hallmark of a quality edit lies in its consistent and unobtrusive adjustment. Modification, when integrated, must be smooth and subtle, preserving the natural flow and aligning seamlessly with the original vision. Therefore, our primary focus is on overcoming the current challe…

2024

FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing

ECCV 2024poster

"Current image editing methods primarily utilize DDIM Inversion, employing a two-branch diffusion approach to preserve the attributes and layout of the original image. However, these methods encounter challenges with non-rigid edits, which involve altering the image’s layout or structure. Our compre…

2024

SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation

AAAI 2024technical

Data augmentation is a crucial component in training neural networks to overcome the limitation imposed by data size, and several techniques have been studied for time series. Although these techniques are effective in certain tasks, they have yet to be generalized to time series benchmarks. We find…

2024

TPC: Test-time Procrustes Calibration for Diffusion-based Human Image Animation

NeurIPS 2024poster

Human image animation aims to generate a human motion video from the inputs of a reference human image and a target motion video. Current diffusion-based image animation systems exhibit high precision in transferring human identity into targeted motion, yet they still exhibit irregular quality in th…

Cited by 3SourcePDFScholar
2024

Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image Editing

ICASSP 2024accepted

In the field of image editing, Null-text Inversion (NTI) enables fine-grained editing while preserving the structure of the original image by optimizing null embeddings during the DDIM sampling process. However, the NTI process is time-consuming, taking more than two minutes per image. To address th…

Cited by 0SourceScholar
2023

Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval

ICASSP 2023accepted

Video Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias…

Cited by 0SourceScholar
2023

ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure

ICLR 2023poster

Studies have shown that modern neural networks tend to be poorly calibrated due to over-confident predictions. Traditionally, post-processing methods have been used to calibrate the model after training. In recent years, various trainable calibration measures have been proposed to incorporate them d…

2023

HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue

EMNLP 2023long findings

Video-grounded Dialogue (VGD) aims to answer questions regarding a given multi-modal input comprising video, audio, and dialogue history. Although there have been numerous efforts in developing VGD systems to improve the quality of their responses, existing systems are competent only to incorporate…

Cited by 0SourcecodeScholar
2023

SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment Retrieval

ICCV 2023poster

Video moment retrieval aims to localize moments in video corresponding to a given language query. To avoid the expensive cost of annotating the temporal moments, weakly-supervised VMR (wsVMR) systems have been studied. For such systems, generating a number of proposals as moment candidates and then…

Cited by 21PDFScholar
2022

Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue

EMNLP 2022main

Video-grounded Dialogue (VGD) aims to decode an answer sentence to a question regarding a given video and dialogue context. Despite the recent success of multi-modal reasoning to generate answer sentences, existing dialogue systems still suffer from a text hallucination problem, which denotes indisc…

2022

SMSMix: Sense-Maintained Sentence Mixup for Word Sense Disambiguation

EMNLP 2022finding

Word Sense Disambiguation (WSD) is an NLP task aimed at determining the correct sense of a word in a sentence from discrete sense choices. Although current systems have attained unprecedented performances for such tasks, the nonuniform distribution of word senses during training generally results in…

2022

Selective Query-Guided Debiasing for Video Corpus Moment Retrieval

ECCV 2022poster

"Video moment retrieval (VMR) aims to localize target moments in untrimmed videos pertinent to a given textual query. Existing retrieval systems tend to rely on retrieval bias as a shortcut and thus, fail to sufficiently learn multi-modal interactions between query and video. This retrieval bias ste…

2021

Structured Co-reference Graph Attention for Video-grounded Dialogue

AAAI 2021technical

A video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given video while keeping track of the dialogue context. Although recent efforts have made great strides in improving the quality…

Cited by 26SourcePDFScholar
2020

VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval

ECCV 2020poster

Video Moment Retrieval (VMR) is a task to localize the temporal moment in untrimmed video specified by natural language query. For VMR, several methods that require full supervision for training have been proposed. Unfortunately, acquiring a large number of training videos with labeled temporal boun…

Cited by 99SourcePDFScholar