← Search

Ji Woo Hong

12 accepted papers

2026

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

ICLR 2026poster

Diffusion Transformers have achieved state-of-the-art performance in class-conditional and multimodal generation, yet the structure of their learned conditional embeddings remains poorly understood. In this work, we present the first systematic study of these embeddings and uncover a notable redunda…

Cited by 0SourceScholar
2026

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding

ICML 2026spotlight

While on-policy distillation offers dense supervision for training small reasoning models, its optimization dynamics in the multimodal domain remain under-explored. In this work, we challenge the standard monolithic view of Vision-Language Model (VLM) distillation by mathematically decomposing the l…

Cited by 0SourceScholar
2026

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signal--rewarding the confidence growth in the ground-truth answer--effectively improves language reasoning training by provi

Cited by 0SourcecodeScholar
2025

FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

ICML 2025spotlight

Drag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts…

Cited by 0SourcePDFScholar
2025

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

CVPR 2025poster

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion Transformer (MDT) for improved handling of both global garment cont…

Cited by 0SourcePDFScholar
2025

Occlusion-robust Stylization for Drawing-based 3D Animation

ICCV 2025poster

3D animation aims to generate a 3D animated video from an input image and a target 3D motion sequence. Recent advances in image-to-3D models enable the creation of animations directly from user-hand drawings. Distinguished from conventional 3D animation, drawing-based 3D animation is crucial to pres…

Cited by 0SourcePDFScholar
2025

TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis

ICCV 2025poster

This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transformers, which offer stable training and continuous transformations for enhanced syn…

Cited by 0SourcePDFScholar
2024

DNI: Dilutional Noise Initialization for Diffusion Video Editing

ECCV 2024poster

"Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving the original structure of the input video. This limitation…

Cited by 2SourcePDFScholar
2024

FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing

ECCV 2024poster

"Current image editing methods primarily utilize DDIM Inversion, employing a two-branch diffusion approach to preserve the attributes and layout of the original image. However, these methods encounter challenges with non-rigid edits, which involve altering the image’s layout or structure. Our compre…

2023

Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval

ICASSP 2023accepted

Video Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias…

Cited by 0SourceScholar
2022

Selective Query-Guided Debiasing for Video Corpus Moment Retrieval

ECCV 2022poster

"Video moment retrieval (VMR) aims to localize target moments in untrimmed videos pertinent to a given textual query. Existing retrieval systems tend to rely on retrieval bias as a shortcut and thus, fail to sufficiently learn multi-modal interactions between query and video. This retrieval bias ste…