← Search

Huiwon Jang

9 accepted papers

2026

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

ICML 2026poster

Augmenting Vision-Language-Action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring …

Cited by 0SourceScholar
2025

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

CVPR 2025poster

Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it would enable the tokenizer to leverage the temporal coherence of videos better for tokenization. Howe…

Cited by 3SourcePDFScholar
2025

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

ICLR 2025oral

Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that on…

2025

Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using Supervised Fine-Tuning (SFT). However, SFT datasets are ofte…

Cited by 0SourceScholar
2024

Adversarial Robustification via Text-to-Image Diffusion Models

ECCV 2024oral

"Adversarial robustness has been conventionally believed as a challenging property to encode for neural networks, requiring plenty of training data. In the recent paradigm of adopting off-the-shelf models, however, access to their training data is often infeasible or not practical, while most of suc…

2024

TrackIME: Enhanced Video Point Tracking via Instance Motion Estimation

NeurIPS 2024spotlight

Tracking points in video frames is essential for understanding video content. However, the task is fundamentally hindered by the computation demands for brute-force correspondence matching across the frames. As the current models down-sample the frame resolutions to mitigate this challenge, they fal…

Cited by 0SourcePDFScholar
2024

Visual Representation Learning with Stochastic Frame Prediction

ICML 2024poster

Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge. This is because of the under-determined nature of frame prediction; multiple potential futures can arise from a single current frame. To tackle this challenge, in thi…

Cited by 3SourcePDFScholar
2023

Modality-Agnostic Self-Supervised Learning with Meta-Learned Masked Auto-Encoder

NeurIPS 2023poster

Despite its practical importance across a wide range of modalities, recent advances in self-supervised learning (SSL) have been primarily focused on a few well-curated domains, e.g., vision and language, often relying on their domain-specific knowledge. For example, Masked Auto-Encoder (MAE) has bec…

Cited by 2SourcePDFScholar
2023

Unsupervised Meta-learning via Few-shot Pseudo-supervised Contrastive Learning

ICLR 2023top-25%

Unsupervised meta-learning aims to learn generalizable knowledge across a distribution of tasks constructed from unlabeled data. Here, the main challenge is how to construct diverse tasks for meta-learning without label information; recent works have proposed to create, e.g., pseudo-labeling via pre…