← Search

Dongyoung Kim

12 accepted papers

2026

Contrastive Representation Regularization for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and pr…

Cited by 0SourceScholar
2026

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

ICML 2026poster

Augmenting Vision-Language-Action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring …

Cited by 0SourceScholar
2026

Verifier-free Test-Time Sampling for Vision Language Action Models

ICLR 2026poster

Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, they remain fundamentally limited in tasks that require high precision due to their single-inference paradigm. While test-time scaling approaches using external verifiers have shown promise, they…

Cited by 0SourceScholar
2025

CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color Constancy

ICCV 2025poster

Computational color constancy, or white balancing, is a key module in a camera's image signal processor (ISP) that corrects color casts from scene lighting. Because this operation occurs in the camera-specific raw color space, white balance algorithms must adapt to different cameras. This paper intr…

Cited by 0SourcePDFScholar
2025

Debiasing Online Preference Learning via Preference Feature Preservation

ACL 2025finding

Recent preference learning frameworks for large language models (LLMs) simplify human preferences with binary pairwise comparisons and scalar rewards. This simplification could make LLMs’ responses biased to mostly preferred features, and would be exacerbated during the iterations of online preferen…

2025

Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using Supervised Fine-Tuning (SFT). However, SFT datasets are ofte…

Cited by 0SourceScholar
2025

Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

ICLR 2025oral

Aligning large language models (LLMs) with human preferences becomes a key component to obtaining state-of-the-art performance, but it yields a huge cost to construct a large human-annotated preference dataset. To tackle this problem, we propose a new framework, Spread Preference Annotation with dir…

Cited by 5SourcePDFScholar
2024

Attentive Illumination Decomposition Model for Multi-Illuminant White Balancing

CVPR 2024poster

White balance (WB) algorithms in many commercial cameras assume single and uniform illumination leading to undesirable results when multiple lighting sources with different chromaticities exist in the scene. Prior research on multi-illuminant WB typically predicts illumination at the pixel level wit…

Cited by 6SourcePDFScholar
2024

Visual Representation Learning with Stochastic Frame Prediction

ICML 2024poster

Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge. This is because of the under-determined nature of frame prediction; multiple potential futures can arise from a single current frame. To tackle this challenge, in thi…

Cited by 3SourcePDFScholar
2023

Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration

NeurIPS 2023poster

A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to struggle in a supervised setup with a task reward, where an…

Cited by 22SourcePDFScholar
2021

Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm Under Mixed Illumination

ICCV 2021poster

We introduce a Large Scale Multi-Illuminant (LSMI) Dataset that contains 7,486 images, captured with three different cameras on more than 2,700 scenes with two or three illuminants. For each image in the dataset, the new dataset provides not only the pixel-wise ground truth illumination but also the…

Cited by 31PDFcodeScholar