← Search

Juno Kim

15 accepted papers

2026

Coverage Improvement and Fast Convergence of On-policy Preference Learning

ICML 2026poster

On-policy preference learning algorithms for language model alignment such as online direct policy optimization (DPO) can significantly outperform their offline counterparts. We provide a theoretical explanation for this phenomenon by analyzing how the sampling policy's coverage evolves throughout o…

Cited by 0SourceScholar
2026

Mirror Mean-Field Langevin Dynamics

ICML 2026poster

The mean-field Langevin dynamics (MFLD) minimizes an entropy-regularized nonlinear convex functional on the Wasserstein space over $\mathbb{R}^d$, and has gained attention recently as a model for the gradient descent dynamics of interacting particle systems such as infinite-width two-layer neural ne…

Cited by 0SourceScholar
2025

CDIS : Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging

IROS 2025

Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previously unseen objects for reliable manipulation and navigation. Existing approaches typically project per-frame 2D instance masks into 3D and merge them, which often

Cited by 0SourceScholar
2025

DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation

ICRA 2025

In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Tasks such as bin-picking and shelfpicking require robust perception to handle occlusions, varying object shapes, and complex spatial arrangements. Traditional RGB

Cited by 2SourceScholar
2025

Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points

NeurIPS 2025poster

Wasserstein gradient flow (WGF) is a common method to perform optimization over the space of probability measures. While WGF is guaranteed to converge to a first-order stationary point, for nonconvex functionals the converged solution does not necessarily satisfy the second-order optimality conditio…

Cited by 0SourceScholar
2025

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

ICML 2025poster

A key paradigm to improve the reasoning capabilities of large language models (LLMs) is to allocate more inference-time compute to search against a verifier or reward model. This process can then be utilized to refine the pretrained model or distill its reasoning patterns into more efficient models.…

Cited by 3SourcePDFScholar
2025

Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression

ICLR 2025poster

We provide a convergence analysis of \emph{deep feature instrumental variable} (DFIV) regression (Xu et al., 2021), a nonparametric approach to IV regression using data-adaptive features learned by deep neural networks in two stages. We prove that the DFIV algorithm achieves the minimax optimal lear…

Cited by 1SourcePDFScholar
2024

$t^3$-Variational Autoencoder: Learning Heavy-tailed Data with Student's t and Power Divergence

ICLR 2024poster

The variational autoencoder (VAE) typically employs a standard normal prior as a regularizer for the probabilistic latent encoder. However, the Gaussian tail often decays too quickly to effectively accommodate the encoded points, failing to preserve crucial structures hidden in the data. In this pap…

Cited by 2SourcePDFScholar
2024

Seg2Grasp: A Robust Modular Suction Grasping in Bin Picking

IROS 2024poster

Current bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in unstructured environments. To overcome these limitations, we introduce Seg2Grasp, a modular pipeline designed for robust suction grasping in dynamic and cluttered b…

Cited by 0SourceScholar
2024

Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems

ICLR 2024spotlight

In this paper, we extend mean-field Langevin dynamics to minimax optimization over probability distributions for the first time with symmetric and provably convergent updates. We propose \emph{mean-field Langevin averaged gradient} (MFL-AG), a single-loop algorithm that implements gradient descent a…

Cited by 10SourcePDFScholar
2024

Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape

ICML 2024oral

Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics of a single layer of attention trained on linear regression tasks. In this paper,…

Cited by 29SourcePDFScholar