← Search

Jaeyeon Kim

17 accepted papers

2026

Any-Order Flexible Length Masked Diffusion

ICLR 2026poster

Masked diffusion models (MDMs) have recently emerged as a promising alternative to autoregressive models over discrete domains. MDMs generate sequences in an any-order, parallel fashion, enabling fast inference and strong performance on non-causal tasks. However, a crucial limitation is that they do…

Cited by 0SourcecodeScholar
2026

Delta Rectified Flow Sampling for Text-to-Image Editing

CVPR 2026

We propose Delta Rectified Flow Sampling (DRFS), a novel inversion-free, path-aware editing framework within rectified flow models for text-to-image editing. DRFS is a distillation-based method that explicitly models the discrepancy between the source and target velocity fields in order to mitigate

Cited by 0SourcecodeScholar
2026

ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models

ICLR 2026poster

Efficient processing of high-resolution images is crucial for real-world vision–language applications. However, existing Large Vision-Language Models (LVLMs) incur substantial computational overhead due to the large number of vision tokens. With the advent of "thinking with images" models, reasoning…

Cited by 0SourcecodeScholar
2026

Fine-Tuning Masked Diffusion for Provable Self-Correction

ICML 2026poster

A natural desideratum for generative models is \emph{self-correction}--detecting and revising low-quality tokens at inference. While Masked Diffusion Models (MDMs) have emerged as a promising approach for generative modeling in discrete spaces, their capacity for self-correction remains poorly under…

Cited by 0SourceScholar
2026

Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning

ICASSP 2026oral

We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over di…

Cited by 0SourcePDFScholar
2026

Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training

ICML 2026poster

Masked Diffusion Models (MDMs) have emerged as a promising approach for generative modeling in discrete spaces. By generating sequences in any order and allowing for parallel decoding, they enable fast inference and strong performance on non-causal tasks. However, this flexibility comes with a *trai…

Cited by 0SourceScholar
2025

Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span

NeurIPS 2025spotlight

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily on motion and contact-based interaction, forecasting human visual perception it…

Cited by 0SourceScholar
2025

LoRA Training Provably Converges to a Low-Rank Global Minimum Or It Fails Loudly (But it Probably Won't Fail)

ICML 2025oral

Low-rank adaptation (LoRA) has become a standard approach for fine-tuning large foundation models. However, our theoretical understanding of LoRA remains limited as prior analyses of LoRA's training dynamics either rely on linearization arguments or consider highly simplified setups. In this work, w…

Cited by 1SourcePDFScholar
2025

Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

ICML 2025oral

In recent years, masked diffusion models (MDMs) have emerged as a promising alternative approach for generative modeling over discrete domains. Compared to autoregressive models (ARMs), MDMs trade off complexity at training time with flexibility at inference time. At training time, they must learn t…

Cited by 4SourcePDFScholar
2024

EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning

ICASSP 2024accepted

We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked codec modeling that improves acoustic awareness of the pret…

Cited by 0SourceScholar
2024

Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset Updates

CVPR 2024poster

Neural radiance field (NeRF) is an emerging technique for 3D scene reconstruction and modeling. However current NeRF-based methods are limited in the capabilities of adding or removing objects. This paper fills the aforementioned gap by proposing a new language-driven method for object manipulation…

2024

Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations

ICASSP 2024accepted

We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures two types of representations with different time resolutions.…

Cited by 0SourceScholar
2024

Optimal Acceleration for Minimax and Fixed-Point Problems is Not Unique

ICML 2024spotlight

Recently, accelerated algorithms using the anchoring mechanism for minimax optimization and fixed-point problems have been proposed, and matching complexity lower bounds establish their optimality. In this work, we present the surprising observation that the optimal acceleration mechanism in minimax…

Cited by 6SourcePDFScholar
2023

Time-Reversed Dissipation Induces Duality Between Minimizing Gradient Norm and Function Value

NeurIPS 2023poster

In convex optimization, first-order optimization methods efficiently minimizing function values have been a central subject study since Nesterov's seminal work of 1983. Recently, however, Kim and Fessler's OGM-G and Lee et al.'s FISTA-G have been presented as alternatives that efficiently minimize t…

Cited by 17SourcePDFScholar