← Search

Chang D. Yoo

60 accepted papers

2026

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

ICLR 2026poster

Diffusion Transformers have achieved state-of-the-art performance in class-conditional and multimodal generation, yet the structure of their learned conditional embeddings remains poorly understood. In this work, we present the first systematic study of these embeddings and uncover a notable redunda…

Cited by 0SourceScholar
2026

Diffusion Negative Preference Optimization Made Simple

ICLR 2026poster

Classifier-Free Guidance (CFG) improves diffusion sampling by encouraging conditional generations while discouraging unconditional ones. Existing preference alignment methods, however, focus only on positive preference pairs, limiting their ability to actively suppress undesirable outputs. Diffusion…

Cited by 0SourcecodeScholar
2026

One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning

ICLR 2026poster

Diffusion Q-Learning (DQL) has established diffusion policies as a high-performing paradigm for offline reinforcement learning, but its reliance on multi-step denoising for action generation renders both training and inference slow and fragile. Existing efforts to accelerate DQL toward one-step deno…

Cited by 0SourceScholar
2026

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signal--rewarding the confidence growth in the ground-truth answer--effectively improves language reasoning training by provi

Cited by 0SourcecodeScholar
2026

TESSAR: Geometry-Aware Active Regression via Dynamic Voronoi Tessellation

ICLR 2026poster

Active learning improves training efficiency by selectively querying the most informative samples for labeling. While it naturally fits classification tasks–where informative samples tend to lie near the decision boundary–its application to regression is less straightforward, as information is distr…

Cited by 0SourceScholar
2026

TOWARDS ROBUST DYSARTHRIC SPEECH RECOGNITION: LLM-AGENT POST-ASR CORRECTION BEYOND WER

ICASSP 2026poster

While Automatic Speech Recognition (ASR) is typically benchmarked by word error rate (WER), real-world applications ultimately hinge on semantic fidelity. This mismatch is particularly problematic for dysarthric speech, where articulatory imprecision and disfluencies can cause severe semantic distor…

Cited by 0SourcePDFScholar
2025

A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion Models

NeurIPS 2025poster

Direct Preference Optimization (DPO) is a key framework for aligning text-to-image models with human preferences, extended by Stepwise Preference Optimization (SPO) to leverage intermediate steps for preference learning, generating more aesthetically pleasing images with significantly less computati…

Cited by 0SourcecodeScholar
2025

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

ICLR 2025poster

In the broader context of deep learning, Multimodal Large Language Models have achieved significant breakthroughs by leveraging powerful Large Language Models as a backbone to align different modalities into the language space. A prime exemplification is the development of Video Large Language Model…

Cited by 0SourcePDFScholar
2025

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

ICML 2025poster

We introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy's confidence, without requiring any auxiliary models or compute. Unlike prior Direct Alignment Algorithms (DAAs) such as…

2025

Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models

ICML 2025poster

Designing effective reward functions remains a fundamental challenge in reinforcement learning (RL), as it often requires extensive human effort and domain expertise. While RL from human feedback has been successful in aligning agents with human intent, acquiring high-quality feedback is costly and…

2025

FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

ICML 2025spotlight

Drag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts…

Cited by 0SourcePDFScholar
2025

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

CVPR 2025poster

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion Transformer (MDT) for improved handling of both global garment cont…

Cited by 0SourcePDFScholar
2025

MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation

ICLR 2025poster

We introduce MDSGen, a novel framework for vision-guided open-domain sound generation optimized for model parameter size, memory consumption, and inference speed. This framework incorporates two key innovations: (1) a redundant video feature removal module that filters out unnecessary visual informa…

Cited by 3SourcePDFScholar
2025

Occlusion-robust Stylization for Drawing-based 3D Animation

ICCV 2025poster

3D animation aims to generate a 3D animated video from an input image and a target 3D motion sequence. Recent advances in image-to-3D models enable the creation of animations directly from user-hand drawings. Distinguished from conventional 3D animation, drawing-based 3D animation is crucial to pres…

Cited by 0SourcePDFScholar
2025

Policy Learning from Large Vision-Language Model Feedback Without Reward Modeling

IROS 2025

Offline reinforcement learning (RL) provides a powerful framework for training robotic agents using pre-collected, suboptimal datasets, eliminating the need for costly, time-consuming, and potentially hazardous online interactions. This is particularly useful in safety-critical real-world applicatio

Cited by 3SourceScholar
2025

Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning

ICASSP 2025accepted

In offline reinforcement learning (RL), learning from fixed datasets presents a promising solution for domains where real-time interaction with the environment is expensive or risky. However, designing dense reward signals for offline dataset requires significant human effort and domain expertise. R…

Cited by 0SourceScholar
2025

Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation

ICASSP 2025accepted

Recent research highlights the potential of multi-modal foundation models in tackling complex decision-making challenges. However, their large parameters make real-world deployment resource-intensive and often impractical for constrained systems. Reinforcement learning (RL) shows promise for task-sp…

Cited by 0SourceScholar
2025

TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis

ICCV 2025poster

This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transformers, which offer stable training and continuous transformations for enhanced syn…

Cited by 0SourcePDFScholar
2024

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

ICASSP 2024accepted

In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapp…

Cited by 0SourceScholar
2024

C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

ICLR 2024poster

In deep learning, test-time adaptation has gained attention as a method for model fine-tuning without the need for labeled data. A prime exemplification is the recently proposed test-time prompt tuning for large-scale vision-language models such as CLIP. Unfortunately, these prompts have been mainly…

2024

Cross-view Masked Diffusion Transformers for Person Image Synthesis

ICML 2024poster

We present X-MDPT ($\underline{Cross}$-view $\underline{M}$asked $\underline{D}$iffusion $\underline{P}$rediction $\underline{T}$ransformers), a novel diffusion model designed for pose-guided human image generation. X-MDPT distinguishes itself by employing masked diffusion transformers that operate…

2024

FRAG: Frequency Adapting Group for Diffusion Video Editing

ICML 2024poster

In video editing, the hallmark of a quality edit lies in its consistent and unobtrusive adjustment. Modification, when integrated, must be smooth and subtle, preserving the natural flow and aligning seamlessly with the original vision. Therefore, our primary focus is on overcoming the current challe…

2024

Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization

IROS 2024poster

Recent studies reveal that well-performing reinforcement learning (RL) agents in training often lack resilience against adversarial perturbations during deployment. This highlights the importance of building a robust agent before deploying it in the real world. Most prior works focus on developing r…

Cited by 0SourcecodeScholar
2024

Progressive Fourier Neural Representation for Sequential Video Compilation

ICLR 2024poster

Neural Implicit Representation (NIR) has recently gained significant attention due to its remarkable ability to encode complex and high-dimensional data into representation space and easily reconstruct it through a trainable mapping function. However, NIR methods assume a one-to-one mapping between…

Cited by 2SourcePDFScholar
2024

Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM

EMNLP 2024finding

The Transformer’s quadratic complexity with input length imposes an unsustainable computational load on large language models (LLMs). In contrast, the Selective Scan Structured State-Space Model, or Mamba, addresses this computational challenge effectively. This paper explores a query-based cross-mo…

Cited by 0SourcePDFScholar
2024

Querying Easily Flip-flopped Samples for Deep Active Learning

ICLR 2024poster

Active learning, a paradigm within machine learning, aims to select and query unlabeled data to enhance model performance strategically. A crucial selection strategy leverages the model's predictive uncertainty, reflecting the informativeness of a data point. While the sample's distance to the decis…

2024

SimPSI: A Simple Strategy to Preserve Spectral Information in Time Series Data Augmentation

AAAI 2024technical

Data augmentation is a crucial component in training neural networks to overcome the limitation imposed by data size, and several techniques have been studied for time series. Although these techniques are effective in certain tasks, they have yet to be generalized to time series benchmarks. We find…

2024

TPC: Test-time Procrustes Calibration for Diffusion-based Human Image Animation

NeurIPS 2024poster

Human image animation aims to generate a human motion video from the inputs of a reference human image and a target motion video. Current diffusion-based image animation systems exhibit high precision in transferring human identity into targeted motion, yet they still exhibit irregular quality in th…

Cited by 3SourcePDFScholar
2024

Unsupervised Speech Recognition with N-skipgram and Positional Unigram Matching

ICASSP 2024accepted

Training unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the power of lower-order N-skipgrams…

Cited by 0SourceScholar
2024

Wavelet-Guided Acceleration of Text Inversion in Diffusion-Based Image Editing

ICASSP 2024accepted

In the field of image editing, Null-text Inversion (NTI) enables fine-grained editing while preserving the structure of the original image by optimizing null embeddings during the DDIM sampling process. However, the NTI process is time-consuming, taking more than two minutes per image. To address th…

Cited by 0SourceScholar
2023

Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval

ICASSP 2023accepted

Video Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias…

Cited by 0SourceScholar
2023

ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure

ICLR 2023poster

Studies have shown that modern neural networks tend to be poorly calibrated due to over-confident predictions. Traditionally, post-processing methods have been used to calibrate the model after training. In recent years, various trainable calibration measures have been proposed to incorporate them d…

2023

Efficient Latent Variable Modeling for Knowledge-Grounded Dialogue Generation

EMNLP 2023long findings

Knowledge-grounded dialogue generation requires first retrieving appropriate external knowledge based on a conversational context and then generating a response grounded on the retrieved knowledge. In general, these two sequential modules, a knowledge retriever and a response generator, have been se…

Cited by 0SourceScholar
2023

HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue

EMNLP 2023long findings

Video-grounded Dialogue (VGD) aims to answer questions regarding a given multi-modal input comprising video, audio, and dialogue history. Although there have been numerous efforts in developing VGD systems to improve the quality of their responses, existing systems are competent only to incorporate…

Cited by 0SourcecodeScholar
2023

On the Soft-Subnetwork for Few-Shot Class Incremental Learning

ICLR 2023poster

Inspired by Regularized Lottery Ticket Hypothesis, which states that competitive smooth (non-binary) subnetworks exist within a dense network, we propose a few-shot class-incremental learning method referred to as Soft-SubNetworks (SoftNet). Our objective is to learn a sequence of sessions increment…

2023

SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment Retrieval

ICCV 2023poster

Video moment retrieval aims to localize moments in video corresponding to a given language query. To avoid the expensive cost of annotating the temporal moments, weakly-supervised VMR (wsVMR) systems have been studied. For such systems, generating a number of proposals as moment candidates and then…

Cited by 21PDFScholar
2022

Decoupled Adversarial Contrastive Learning for Self-Supervised Adversarial Robustness

ECCV 2022poster

"\textit{Adversarial training} (AT) for robust representation learning and \textit{self-supervised learning} (SSL) for unsupervised representation learning are two active research fields. Integrating AT into SSL, multiple prior works have accomplished a highly significant yet challenging task: learn…

2022

Dual Temperature Helps Contrastive Learning Without Many Negative Samples: Towards Understanding and Simplifying MoCo

CVPR 2022poster

Contrastive learning (CL) is widely known to require many negative samples, 65536 in MoCo for instance, for which the performance of a dictionary-free framework is often inferior because the negative sample size (NSS) is limited by its mini-batch size (MBS). To decouple the NSS from the MBS, a dynam…

Cited by 59PDFcodeScholar
2022

Fast and Efficient MMD-Based Fair PCA via Optimization over Stiefel Manifold

AAAI 2022technical

This paper defines fair principal component analysis (PCA) as minimizing the maximum mean discrepancy (MMD) between the dimensionality-reduced conditional distributions of different protected classes. The incorporation of MMD naturally leads to an exact and tractable mathematical formulation of fair…

2022

Forget-free Continual Learning with Winning Subnetworks

ICML 2022spotlight

Inspired by Lottery Ticket Hypothesis that competitive subnetworks exist within a dense network, we propose a continual learning method referred to as Winning SubNetworks (WSN), which sequentially learns and selects an optimal subnetwork for each task. Specifically, WSN jointly learns the model weig…

2022

How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning

ICLR 2022poster

To avoid collapse in self-supervised learning (SSL), a contrastive loss is widely used but often requires a large number of negative samples. Without negative samples yet achieving competitive performance, a recent work~\citep{chen2021exploring} has attracted significant attention for providing a mi…

Cited by 98SourcePDFScholar
2022

Selective Query-Guided Debiasing for Video Corpus Moment Retrieval

ECCV 2022poster

"Video moment retrieval (VMR) aims to localize target moments in untrimmed videos pertinent to a given textual query. Existing retrieval systems tend to rely on retrieval bias as a shortcut and thus, fail to sufficiently learn multi-modal interactions between query and video. This retrieval bias ste…

2022

SoftGroup for 3D Instance Segmentation on Point Clouds

CVPR 2022oral

Existing state-of-the-art 3D instance segmentation methods perform semantic segmentation followed by grouping. The hard predictions are made when performing semantic segmentation such that each point is associated with a single class. However, the errors stemming from hard decision propagate into gr…

Cited by 297PDFcodeScholar
2021

Robust Maml: Prioritization Task Buffer with Adaptive Learning Process for Model-Agnostic Meta-Learning

ICASSP 2021accepted

Model agnostic meta-learning (MAML) is a popular state-of-the-art meta-learning algorithm that provides good weight initialization of a model given a variety of learning tasks. The model initialized by provided weight can be fine-tuned to an unseen task despite only using a small amount of samples a…

Cited by 0SourceScholar
2021

SCNet: Training Inference Sample Consistency for Instance Segmentation

AAAI 2021technical

Cascaded architectures have brought significant performance improvement in object detection and instance segmentation. However, there are lingering issues regarding the disparity in the Intersection-over-Union (IoU) distribution of the samples between training and inference. This disparity can poten…

2021

Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model

IROS 2021poster

Developing an agent in reinforcement learning (RL) that is capable of performing complex control tasks directly from high-dimensional observation such as raw pixels is a challenge as efforts still need to be made towards improving sample efficiency and generalization of RL algorithm. This paper cons…

Cited by 24SourceScholar
2021

Structured Co-reference Graph Attention for Video-grounded Dialogue

AAAI 2021technical

A video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given video while keeping track of the dialogue context. Although recent efforts have made great strides in improving the quality…

Cited by 26SourcePDFScholar
2020

Modality Shifting Attention Network for Multi-Modal Video Question Answering

CVPR 2020poster

This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on…

Cited by 102PDFScholar
2020

VLANet: Video-Language Alignment Network for Weakly-Supervised Video Moment Retrieval

ECCV 2020poster

Video Moment Retrieval (VMR) is a task to localize the temporal moment in untrimmed video specified by natural language query. For VMR, several methods that require full supervision for training have been proposed. Unfortunately, acquiring a large number of training videos with labeled temporal boun…

Cited by 99SourcePDFScholar
2019

Progressive Attention Memory Network for Movie Story Question Answering

CVPR 2019poster

This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour,…

Cited by 96PDFScholar
2018

Pivot Correlational Neural Network for Multimodal Video Categorization

ECCV 2018poster

This paper considers an architecture for multimodal video categorization referred to as Pivot Correlational Neural Network (Pivot CorrNN). The architecture is trained to maximizes the correlation between the hidden states as well as the predictions of the modal-agnostic pivot stream and modal-specif…

Cited by 14SourcePDFScholar
2015

Dense Image Registration and Deformable Surface Reconstruction in Presence of Occlusions and Minimal Texture

ICCV 2015poster

Deformable surface tracking from monocular images is well-known to be under-constrained. Occlusions often make the task even more challenging, and can result in failure if the surface is not sufficiently textured. In this work, we explicitly address the problem of 3D reconstruction of poorly texture…

Cited by 53PDFScholar