← Search

Jinwoo Shin

158 accepted papers

2026

Contrastive Representation Regularization for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and pr…

Cited by 0SourceScholar
2026

Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling

ICLR 2026poster

Denoising generative models, such as diffusion and flow-based models, produce high-quality samples but require many denoising steps due to discretization error. Flow maps, which estimate the average velocity between timesteps, mitigate this error and enable faster sampling. However, their training t…

Cited by 0SourcecodeScholar
2026

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

ICML 2026poster

Augmenting Vision-Language-Action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring …

Cited by 0SourceScholar
2026

HAMLET: Switch Your Vision-Language-Action Model into a History-Aware Policy

ICLR 2026poster

Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Language-Action models (VLAs) have been designed without considering this aspect, i.e., they rely solely on the current observation, ignoring preceding context. In…

Cited by 0SourceScholar
2026

Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance

CVPR 2026

Recent text-to-video (T2V) models have demonstrated strong capabilities in producing high-quality, dynamic videos. To improve the visual controllability, recent works have considered fine-tuning pre-trained T2V models to support image-to-video (I2V) generation. However, such adaptation frequently su

Cited by 0SourcecodeScholar
2026

Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation

CVPR 2026

Achieving precise alignment between user intent and generated visuals remains a central challenge in text-to-visual generation, as a single attempt often fails to produce the desired output. To handle this, prior approaches mainly scale the visual generation process (e.g., increasing sampling steps

Cited by 0SourceScholar
2026

Verifier-free Test-Time Sampling for Vision Language Action Models

ICLR 2026poster

Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, they remain fundamentally limited in tasks that require high precision due to their single-inference paradigm. While test-time scaling approaches using external verifiers have shown promise, they…

Cited by 0SourceScholar
2026

Vision-aligned Latent Reasoning for Multi-Modal Large Language Model

ICML 2026poster

Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems which require extensive multi-step reasoning. This is primarily due to the progressive dilution of visual information during long-context generation, which…

Cited by 0SourceScholar
2025

Accelerated Test-Time Scaling with Model-Free Speculative Sampling

EMNLP 2025

Language models have demonstrated remarkable capabilities in reasoning tasks through test-time scaling techniques like best-of-N sampling and tree search. However, these approaches often demand substantial computational resources, creating a critical trade-off between performance and efficiency. We

Cited by 0SourcePDFScholar
2025

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

CVPR 2025poster

Aligning text-to-image (T2I) diffusion models with prefer-ence optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting…

Cited by 5SourcePDFScholar
2025

Closest Neighbors are Harmful for Lightweight Masked Auto-encoders

CVPR 2025poster

Learning the visual representation via masked auto-encoder (MAE) training has been proven to be a powerful technique. Transferring the pre-trained vision transformer (ViT) to downstream tasks leads to superior performance compared to conventional task-by-task supervised learning. Recent research wo…

2025

Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers

NeurIPS 2025poster

As large language models increasingly gain popularity in real-world applications, processing extremely long contexts, often exceeding the model’s pre-trained context limits, has emerged as a critical challenge. While existing approaches to efficient long-context processing show promise, recurrent co…

Cited by 0SourceScholar
2025

Controllable Blur Data Augmentation Using 3D-Aware Motion Estimation

ICLR 2025poster

Existing realistic blur datasets provide insufficient variety in scenes and blur patterns to be trained, while expanding data diversity demands considerable time and effort due to complex dual-camera systems. To address the challenge, data augmentation can be an effective way to artificially increas…

Cited by 0SourcePDFScholar
2025

Controllable Human Image Generation with Personalized Multi-Garments

CVPR 2025poster

We present BootControl, a novel framework based on text-to-image diffusion models for controllable human image generation with multiple reference garments.Here, the main bottleneck is data acquisition for training: collecting a large-scale dataset of high-quality reference garment images per human s…

Cited by 0SourcePDFScholar
2025

Debiasing Online Preference Learning via Preference Feature Preservation

ACL 2025finding

Recent preference learning frameworks for large language models (LLMs) simplify human preferences with binary pairwise comparisons and scalar rewards. This simplification could make LLMs’ responses biased to mostly preferred features, and would be exacerbated during the iterations of online preferen…

2025

DiffusionGuard: A Robust Defense Against Malicious Diffusion-based Image Editing

ICLR 2025poster

Recent advances in diffusion models have introduced a new era of text-guided image manipulation, enabling users to create realistic edited images with simple textual prompts. However, there is significant concern about the potential misuse of these methods, especially in creating misleading or harmf…

2025

Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction

CVPR 2025poster

Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it would enable the tokenizer to leverage the temporal coherence of videos better for tokenization. Howe…

Cited by 3SourcePDFScholar
2025

Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents

ICLR 2025poster

Recent advances in large language models (LLMs) have led to a growing interest in developing LLM-based agents for automating web tasks. However, these agents often struggle with even simple tasks on real-world websites due to their limited capability to understand and process complex web page struct…

2025

MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement

NeurIPS 2025poster

Agents based on large language models (LLMs) for machine learning engineering (MLE) can automatically implement ML models via code generation. However, existing approaches to build such agents often rely heavily on inherent LLM knowledge and employ coarse exploration strategies that modify the entir…

Cited by 0SourceScholar
2025

Mamba Drafters for Speculative Decoding

EMNLP 2025

Speculative decoding has emerged as a promising approach to accelerating large language model (LLM) generation using a fast drafter while maintaining alignment with the target model’s distribution. However, existing approaches face a trade-off: external drafters offer flexibility but can suffer from

2025

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

ICML 2025poster

Selecting a layer normalization (LN) strategy that stabilizes training and speeds convergence in Transformers remains difficult, even for today’s large language models (LLM). We present a comprehensive analytical foundation for understanding how different LN strategies influence training dynamics in…

Cited by 0SourcePDFScholar
2025

Personalized Language Models via Privacy-Preserving Evolutionary Model Merging

EMNLP 2025

Personalization in language models aims to tailor model behavior to individual users or user groups. Prompt-based methods incorporate user preferences into queries, while training-based methods encode them into model parameters. Model merging has also been explored for personalization under limited

2025

ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification

ICML 2025poster

Self-awareness, i.e., the ability to assess and correct one's generation, is a fundamental aspect of human intelligence, making its replication in large language models (LLMs) an important yet challenging task. Previous works tackle this by employing extensive reinforcement learning or relying on la…

Cited by 3SourcePDFScholar
2025

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

ICLR 2025oral

Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that on…

2025

Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using Supervised Fine-Tuning (SFT). However, SFT datasets are ofte…

Cited by 0SourceScholar
2025

Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

ICLR 2025oral

Aligning large language models (LLMs) with human preferences becomes a key component to obtaining state-of-the-art performance, but it yields a huge cost to construct a large human-annotated preference dataset. To tackle this problem, we propose a new framework, Spread Preference Annotation with dir…

Cited by 5SourcePDFScholar
2025

StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment

IJCAI 2025

Learning robust representations from data often requires scale, which has led to the success of recent zero-shot models such as CLIP. However, the obtained robustness can easily be deteriorated when these models are fine-tuned on other downstream tasks (e.g., of smaller scales). Previous works often

2025

Subtask-Aware Visual Reward Learning from Segmented Demonstrations

ICLR 2025poster

Reinforcement Learning (RL) agents have demonstrated their potential across various robotic tasks. However, they still heavily rely on human-engineered reward functions, requiring extensive trial-and-error and access to target behavior information, often unavailable in real-world settings. This pape…

Cited by 0SourcePDFScholar
2025

Test-Time Adaptation with Binary Feedback

ICML 2025poster

Deep learning models perform poorly when domain shifts exist between training and test data. Test-time adaptation (TTA) is a paradigm to mitigate this issue by adapting pre-trained models using only unlabeled test samples. However, existing TTA methods can fail under severe domain shifts, while rece…

2025

Think Clearly: Improving Reasoning via Redundant Token Pruning

EMNLP 2025

Recent large language models have shown promising capabilities in long-form reasoning, following structured chains of thought before arriving at a final answer. However, we observe that these reasoning paths tend to include substantial redundancy; analyzing attention patterns reveals that attention

Cited by 0SourcePDFScholar
2025

Training Text-to-Molecule Models with Context-Aware Tokenization

EMNLP 2025

Recently, text-to-molecule models have shown great potential across various chemical applications, e.g., drug-discovery. These models adapt language models to molecular data by representing molecules as sequences of atoms. However, they rely on atom-level tokenizations, which primarily focus on mode

2024

Conditional Synthesis of 3D Molecules with Time Correction Sampler

NeurIPS 2024poster

Diffusion models have demonstrated remarkable success in various domains, including molecular generation. However, conditional molecular generation remains a fundamental challenge due to an intrinsic trade-off between targeting specific chemical properties and generating meaningful samples from the…

Cited by 2SourcePDFScholar
2024

Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models

ICLR 2024poster

Fine-tuning text-to-image models with reward functions trained on human feedback data has proven effective for aligning model behavior with human intent. However, excessive optimization with such reward models, which serve as mere proxy objectives, can compromise the performance of fine-tuned models…

2024

Data-Efficient Molecular Generation with Hierarchical Textual Inversion

ICML 2024poster

Developing an effective molecular generation framework even with a limited number of molecules is often important for its practical deployment, e.g., drug discovery, since acquiring task-related molecular data requires expensive and time-consuming experimental costs. To tackle this issue, we introdu…

2024

Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models

NeurIPS 2024poster

Text-to-image (T2I) diffusion models, when fine-tuned on a few personal images, can generate visuals with a high degree of consistency. However, such fine-tuned models are not robust; they often fail to compose with concepts of pretrained model or other fine-tuned models. To address this, we propose…

Cited by 1SourcePDFScholar
2024

Discovering and Mitigating Visual Biases through Keyword Explanation

CVPR 2024highlight

Addressing biases in computer vision models is crucial for real-world AI deployments. However mitigating visual biases is challenging due to their unexplainable nature often identified indirectly through visualization or sample statistics which necessitates additional human supervision for interpret…

2024

DreamFlow: High-quality text-to-3D generation by Approximating Probability Flow

ICLR 2024spotlight

Recent progress in text-to-3D generation has been achieved through the utilization of score distillation methods: they make use of the pre-trained text-to-image (T2I) diffusion models by distilling via the diffusion model training objective. However, such an approach inevitably results in the use of…

Cited by 16SourcePDFScholar
2024

Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

ICLR 2024poster

Video diffusion models have recently made great progress in generation quality, but are still limited by the high memory and computational requirements. This is because current video diffusion models often attempt to process high-dimensional videos directly. To tackle this issue, we propose content-…

Cited by 24SourcePDFScholar
2024

Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs

ICLR 2024poster

Large language models (LLMs) have shown remarkable performance in various natural language processing tasks. However, a primary constraint they face is the context limit, i.e., the maximum number of tokens they can process. Previous works have explored architectural changes and modifications in posi…

2024

Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback

EMNLP 2024finding

Large language models (LLMs) fine-tuned with alignment techniques, such as reinforcement learning from human feedback, have been instrumental in developing some of the most capable AI systems to date. Despite their success, existing methods typically rely on simple binary labels, such as those indic…

2024

Online Adaptation of Language Models with a Memory of Amortized Contexts

NeurIPS 2024poster

Due to the rapid generation and dissemination of information, large language models (LLMs) quickly run out of date despite enormous development costs. To address the crucial need to keep models updated, online learning has emerged as a critical tool when utilizing LLMs for real-world applications. H…

2024

Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning

NeurIPS 2024poster

In tabular prediction tasks, tree-based models combined with automated feature engineering methods often outperform deep learning approaches that rely on learned representations. While these feature engineering techniques are effective, they typically depend on a pre-defined search space and primari…

2024

Querying Easily Flip-flopped Samples for Deep Active Learning

ICLR 2024poster

Active learning, a paradigm within machine learning, aims to select and query unlabeled data to enhance model performance strategically. A crucial selection strategy leverages the model's predictive uncertainty, reflecting the informativeness of a data point. While the sample's distance to the decis…

2024

Real-World Efficient Blind Motion Deblurring via Blur Pixel Discretization

CVPR 2024poster

As recent advances in mobile camera technology have enabled the capability to capture high-resolution images such as 4K images the demand for an efficient deblurring model handling large motion has increased. In this paper we discover that the image residual errors i.e. blur-sharp pixel differences…

Cited by 5SourcePDFScholar
2024

Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion

ECCV 2024poster

"This paper addresses the societal concerns arising from large-scale text-to-image diffusion models for generating potentially harmful or copyrighted content. Existing models rely heavily on internet-crawled data, wherein problematic concepts persist due to incomplete filtration processes. While pre…

2024

SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs

ICLR 2024poster

Large language models (LLMs) have made significant advancements in various natural language processing tasks, including question answering (QA) tasks. While incorporating new information with the retrieval of relevant passages is a promising way to improve QA with LLMs, the existing methods often re…

2024

TrackIME: Enhanced Video Point Tracking via Instance Motion Estimation

NeurIPS 2024spotlight

Tracking points in video frames is essential for understanding video content. However, the task is fundamentally hindered by the computation demands for brute-force correspondence matching across the frames. As the current models down-sample the frame resolutions to mitigate this challenge, they fal…

Cited by 0SourcePDFScholar
2024

Visual Representation Learning with Stochastic Frame Prediction

ICML 2024poster

Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge. This is because of the under-determined nature of frame prediction; multiple potential futures can arise from a single current frame. To tackle this challenge, in thi…

Cited by 3SourcePDFScholar
2023

Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration

NeurIPS 2023poster

A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to struggle in a supervised setup with a task reward, where an…

Cited by 22SourcePDFScholar
2023

BiasAdv: Bias-Adversarial Augmentation for Model Debiasing

CVPR 2023poster

Neural networks are often prone to bias toward spurious correlations inherent in a dataset, thus failing to generalize unbiased test criteria. A key challenge to resolving the issue is the significant lack of bias-conflicting training data (i.e., samples without spurious correlations). In this paper…

Cited by 30SourcePDFScholar
2023

Collaborative Score Distillation for Consistent Visual Editing

NeurIPS 2023poster

Generative priors of large-scale text-to-image diffusion models enable a wide range of new generation and editing applications on diverse visual modalities. However, when adapting these priors to complex visual modalities, often represented as multiple images (e.g., video or 3D scene), achieving con…

Cited by 22SourcePDFScholar
2023

Confidence-Aware Training of Smoothed Classifiers for Certified Robustness

AAAI 2023technical

Any classifier can be "smoothed out" under Gaussian noise to build a new classifier that is provably robust to l2-adversarial perturbations, viz., by averaging its predictions over the noise via randomized smoothing. Under the smoothed classifiers, the fundamental trade-off between accuracy and (adv…

2023

Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles

AISTATS 2023poster

We study contextual linear bandit problems under feature uncertainty; they are noisy with missing entries. To address the challenges of the noise, we analyze Bayesian oracles given observed noisy features. Our Bayesian analysis finds that the optimal hypothesis can be far from the underlying realiza…

2023

Enhancing Multiple Reliability Measures via Nuisance-Extended Information Bottleneck

CVPR 2023poster

In practical scenarios where training data is limited, many predictive signals in the data can be rather from some biases in data acquisition (i.e., less generalizable), so that one cannot prevent a model from co-adapting on such (so-called) "shortcut" signals: this makes the model fragile in variou…

2023

Guide Your Agent with Adaptive Multimodal Rewards

NeurIPS 2023poster

Developing an agent capable of adapting to unseen environments remains a difficult challenge in imitation learning. This work presents Adaptive Return-conditioned Policy (ARP), an efficient framework designed to enhance the agent's generalization ability using natural language task descriptions and…

2023

Guiding Energy-based Models via Contrastive Latent Variables

ICLR 2023top-25%

An energy-based model (EBM) is a popular generative framework that offers both explicit density and architectural flexibility, but training them is difficult since it is often unstable and time-consuming. In recent years, various training techniques have been developed, e.g., better divergence measu…

2023

IFSeg: Image-Free Semantic Segmentation via Vision-Language Model

CVPR 2023poster

Vision-language (VL) pre-training has recently gained much attention for its transferability and flexibility in novel concepts (e.g., cross-modality transfer) across various visual tasks. However, VL-driven segmentation has been under-explored, and the existing approaches still have the burden of ac…

2023

Imitating Graph-Based Planning with Goal-Conditioned Policies

ICLR 2023poster

Recently, graph-based planning algorithms have gained much attention to solve goal-conditioned reinforcement learning (RL) tasks: they provide a sequence of subgoals to reach the target-goal, and the agents learn to execute subgoal-conditioned policies. However, the sample-efficiency of such RL sche…

2023

Learning Large-scale Neural Fields via Context Pruned Meta-Learning

NeurIPS 2023poster

We introduce an efficient optimization-based meta-learning technique for large-scale neural field training by realizing significant memory savings through automated online context point selection. This is achieved by focusing each learning step on the subset of data with the highest expected immedia…

2023

Modality-Agnostic Self-Supervised Learning with Meta-Learned Masked Auto-Encoder

NeurIPS 2023poster

Despite its practical importance across a wide range of modalities, recent advances in self-supervised learning (SSL) have been primarily focused on a few well-curated domains, e.g., vision and language, often relying on their domain-specific knowledge. For example, Masked Auto-Encoder (MAE) has bec…

Cited by 2SourcePDFScholar
2023

Modality-Agnostic Variational Compression of Implicit Neural Representations

ICML 2023poster

We introduce a modality-agnostic neural compression algorithm based on a functional view of data and parameterised as an Implicit Neural Representation (INR). Bridging the gap between latent coding and sparsity, we obtain compact latent representations non-linearly mapped to a soft gating mechanism.…

Cited by 27SourcePDFScholar
2023

Multi-View Masked World Models for Visual Robotic Manipulation

ICML 2023poster

Visual robotic manipulation research and applications often use multiple cameras, or views, to better perceive the world. How else can we utilize the richness of multi-view data? In this paper, we investigate how to learn good representations with multi-view data and utilize them for visual robotic…

2023

Prefer to Classify: Improving Text Classifiers via Auxiliary Preference Learning

ICML 2023poster

The development of largely human-annotated benchmarks has driven the success of deep neural networks in various NLP tasks. To enhance the effectiveness of existing benchmarks, collecting new additional input-output pairs is often too costly and challenging, particularly considering their marginal im…

2023

Preference Transformer: Modeling Human Preferences using Transformers for RL

ICLR 2023poster

Preference-based reinforcement learning (RL) provides a framework to train agents using human preferences between two behaviors. However, preference-based RL has been challenging to scale since it requires a large amount of human feedback to learn a reward function aligned with human intent. In this…

2023

S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist Captions

NeurIPS 2023poster

Vision-language models, such as contrastive language-image pre-training (CLIP), have demonstrated impressive results in natural image domains. However, these models often struggle when applied to specialized domains like remote sensing, and adapting to such domains is challenging due to the limited…

2023

STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled Tables

ICLR 2023top-25%

Learning with few labeled tabular samples is often an essential requirement for industrial machine learning applications as varieties of tabular data suffer from high annotation costs or have difficulties in collecting new samples for novel tasks. Despite the utter importance, such a problem is quit…

2023

Slimmed Asymmetrical Contrastive Learning and Cross Distillation for Lightweight Model Training

NeurIPS 2023poster

Contrastive learning (CL) has been widely investigated with various learning mechanisms and achieves strong capability in learning representations of data in a self-supervised manner using unlabeled data. A common fashion of contrastive learning on this line is employing mega-sized encoders to achie…

2023

String-Based Molecule Generation Via Multi-Decoder VAE

ICASSP 2023accepted

In this study, we investigate the problem of string-based molecular generation via variational autoencoders (VAEs) that have served a popular generative approach for various tasks in artificial intelligence. Our main idea is to maintain multiple decoders while sharing a single encoder, i.e., it is a…

Cited by 4SourceScholar
2023

Unsupervised Meta-learning via Few-shot Pseudo-supervised Contrastive Learning

ICLR 2023top-25%

Unsupervised meta-learning aims to learn generalizable knowledge across a distribution of tasks constructed from unlabeled data. Here, the main challenge is how to construct diverse tasks for meta-learning without label information; recent works have proposed to create, e.g., pseudo-labeling via pre…

2023

infoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-information

ACL 2023long

The success of NLP systems often relies on the availability of large, high-quality datasets. However, not all samples in these datasets are equally valuable for learning, as some may be redundant or noisy. Several methods for characterizing datasets based on model-driven meta-information (e.g., mode…

2022

Consistency Regularization for Adversarial Robustness

AAAI 2022technical

Adversarial training (AT) is currently one of the most successful methods to obtain the adversarial robustness of deep neural networks. However, the phenomenon of robust overfitting, i.e., the robustness starts to decrease significantly during AT, has been problematic, not only making practitioners…

2022

Contrastive Dual Gating: Learning Sparse Features With Contrastive Learning

CVPR 2022poster

Contrastive learning (or its variants) has recently become a promising direction in the self-supervised learning domain, achieving similar performance as supervised learning with minimum fine-tuning. Despite the labeling efficiency, wide and large networks are required to achieve high accuracy, whic…

Cited by 14PDFScholar
2022

Disentangling Sources of Risk for Distributional Multi-Agent Reinforcement Learning

ICML 2022spotlight

In cooperative multi-agent reinforcement learning, the outcomes of agent-wise policies are highly stochastic due to the two sources of risk: (a) random actions taken by teammates and (b) random transition and rewards. Although the two sources have very distinct characteristics, existing frameworks a…

Cited by 12SourcePDFScholar
2022

Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

ICLR 2022poster

In the deep learning era, long video generation of high-quality still remains challenging due to the spatio-temporal complexity and continuity of videos. Existing prior works have attempted to model video distribution by representing videos as 3D grids of RGB values, which impedes the scale of gener…

Cited by 234SourcePDFScholar
2022

K-Centered Patch Sampling for Efficient Video Recognition

ECCV 2022poster

"For decades, it has been a common practice to choose a subset of video frames for reducing the computational burden of a video understanding model. In this paper, we argue that this popular heuristic might be sub-optimal under recent transformer-based models. Specifically, inspired by that transfor…

2022

Meta-Learning with Self-Improving Momentum Target

NeurIPS 2022accept

The idea of using a separately trained target model (or teacher) to improve the performance of the student model has been increasingly popular in various machine learning domains, and meta-learning is no exception; a recent discovery shows that utilizing task-wise target models can significantly boo…

2022

NOTE: Robust Continual Test-time Adaptation Against Temporal Correlation

NeurIPS 2022accept

Test-time adaptation (TTA) is an emerging paradigm that addresses distributional shifts between training and testing phases without additional data acquisition or labeling cost; only unlabeled test data streams are used for continual model adaptation. Previous TTA schemes assume that the test sample…

2022

Patch-Level Representation Learning for Self-Supervised Vision Transformers

CVPR 2022oral

Recent self-supervised learning (SSL) methods have shown impressive results in learning visual representations from unlabeled images. This paper aims to improve their performance further by utilizing the architectural advantages of the underlying neural network, as the current state-of-the-art visua…

Cited by 66PDFcodeScholar
2022

SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

ICLR 2022poster

Preference-based reinforcement learning (RL) has shown potential for teaching agents to perform the target tasks without a costly, pre-defined reward function by learning the reward with a supervisor’s preference between the two agent behaviors. However, preference-based learning often requires a la…

Cited by 113SourcePDFScholar
2022

Saliency Grafting: Innocuous Attribution-Guided Mixup with Calibrated Label Mixing

AAAI 2022technical

The Mixup scheme suggests mixing a pair of samples to create an augmented training sample and has gained considerable attention recently for improving the generalizability of neural networks. A straightforward and widely used extension of Mixup is to combine with regional dropout-like methods: remov…

Cited by 25SourcePDFScholar
2022

Scalable Neural Video Representations with Learnable Positional Features

NeurIPS 2022accept

Succinct representation of complex signals using coordinate-based neural representations (CNRs) has seen great progress, and several recent efforts focus on extending them for handling videos. Here, the main challenge is how to (a) alleviate a compute-inefficiency in training CNRs to (b) achieve hig…

2022

Self-Supervised Dense Consistency Regularization for Image-to-Image Translation

CVPR 2022poster

Unsupervised image-to-image translation has gained considerable attention due to the recent impressive progress based on generative adversarial networks (GANs). In this paper, we present a simple but effective regularization technique for improving GAN-based image-to-image translation. To generate i…

Cited by 25PDFScholar
2022

Spread Spurious Attribute: Improving Worst-group Accuracy with Spurious Attribute Estimation

ICLR 2022poster

The paradigm of worst-group loss minimization has shown its promise in avoiding to learn spurious correlations, but requires costly additional supervision on spurious attributes. To resolve this, recent works focus on developing weaker forms of supervision---e.g., hyperparameters discovered with a s…

Cited by 105SourcePDFScholar
2022

TSPipe: Learn from Teacher Faster with Pipelines

ICML 2022spotlight

The teacher-student (TS) framework, training a (student) network by utilizing an auxiliary superior (teacher) network, has been adopted as a popular training paradigm in many machine learning schemes, since the seminal work—Knowledge distillation (KD) for model compression and transfer learning. Man…

Cited by 1SourcePDFScholar
2022

Time Is MattEr: Temporal Self-supervision for Video Transformers

ICML 2022spotlight

Understanding temporal dynamics of video is an essential aspect of learning better video representations. Recently, transformer-based architectural designs have been extensively explored for video tasks due to their capability to capture long-term dependency of input sequences. However, we found tha…

2022

What Makes Better Augmentation Strategies? Augment Difficult but Not too Different

ICLR 2022poster

The practice of data augmentation has been extensively used to boost the performance of deep neural networks for various NLP tasks. It is more effective when only a limited number of labeled samples is available, e.g., low-data or class-imbalanced regimes. Most current augmentation techniques rely o…

Cited by 15SourcePDFScholar
2021

$i$-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning

ICLR 2021poster

Contrastive representation learning has shown to be effective to learn representations from unlabeled data. However, much progress has been made in vision domains relying on data augmentations carefully designed using domain knowledge. In this work, we propose i-Mix, a simple yet effective domain-ag…

2021

GTA: Graph Truncated Attention for Retrosynthesis

AAAI 2021technical

Retrosynthesis is the task of predicting reactant molecules from a given product molecule and is, important in organic chemistry because the identification of a synthetic path is as demanding as the discovery of new chemical compounds. Recently, the retrosynthesis task has been solved automatically…

Cited by 71SourcePDFScholar
2021

Improving Transferability of Representations via Augmentation-Aware Self-Supervision

NeurIPS 2021poster

Recent unsupervised representation learning methods have shown to be effective in a range of vision tasks by learning representations invariant to data augmentations such as random cropping and color jittering. However, such invariance could be harmful to downstream tasks if they rely on the charact…

2021

Landmark-Guided Subgoal Generation in Hierarchical Reinforcement Learning

NeurIPS 2021poster

Goal-conditioned hierarchical reinforcement learning (HRL) has shown promising results for solving complex and long-horizon RL tasks. However, the action space of high-level policy in the goal-conditioned HRL is often large, so it results in poor exploration, leading to inefficiency in training. In…

2021

Layer-adaptive Sparsity for the Magnitude-based Pruning

ICLR 2021poster

Recent discoveries on neural network pruning reveal that, with a carefully chosen layerwise sparsity, a simple magnitude-based pruning achieves state-of-the-art tradeoff between sparsity and performance. However, without a clear consensus on ``how to choose,'' the layerwise sparsities are mostly sel…

2021

Learning to Sample with Local and Global Contexts in Experience Replay Buffer

ICLR 2021poster

Experience replay, which enables the agents to remember and reuse experience from the past, has played a significant role in the success of off-policy reinforcement learning (RL). To utilize the experience replay efficiently, the existing sampling methods allow selecting out more meaningful experien…

Cited by 26SourcePDFScholar
2021

MASKER: Masked Keyword Regularization for Reliable Text Classification

AAAI 2021technical

Pre-trained language models have achieved state-of-the-art accuracies on various text classification tasks, e.g., sentiment analysis, natural language inference, and semantic textual similarity. However, the reliability of the fine-tuned text classifiers is an often underlooked performance criterion…

2021

Object-Aware Regularization for Addressing Causal Confusion in Imitation Learning

NeurIPS 2021poster

Behavioral cloning has proven to be effective for learning sequential decision-making policies from expert demonstrations. However, behavioral cloning often suffers from the causal confusion problem where a policy relies on the noticeable effect of expert actions due to the strong correlation but no…

2021

Object-aware Contrastive Learning for Debiased Scene Representation

NeurIPS 2021poster

Contrastive self-supervised learning has shown impressive results in learning visual representations from unlabeled images by enforcing invariance against different data augmentations. However, the learned representations are often contextually biased to the spurious scene correlations of different…

2021

Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble

CoRL 2021poster

Recent advance in deep offline reinforcement learning (RL) has made it possible to train strong robotic agents from offline datasets. However, depending on the quality of the trained agents and the application being considered, it is often desirable to fine-tune such agents via further online intera…

Cited by 239SourcecodeScholar
2021

Quality-Agnostic Image Recognition via Invertible Decoder

CVPR 2021poster

Despite the remarkable performance of deep models on image recognition tasks, they are known to be susceptible to common corruptions such as blur, noise, and low-resolution. Data augmentation is a conventional way to build a robust model by considering these common corruptions during the training. H…

Cited by 30PDFScholar
2021

RetCL: A Selection-based Approach for Retrosynthesis via Contrastive Learning

IJCAI 2021poster

Retrosynthesis, of which the goal is to find a set of reactants for synthesizing a target product, is an emerging research area of deep learning. While the existing approaches have shown promising results, they currently lack the ability to consider availability (e.g., stability or purchasability) o…

Cited by 23SourcePDFScholar
2021

RoMA: Robust Model Adaptation for Offline Model-based Optimization

NeurIPS 2021poster

We consider the problem of searching an input maximizing a black-box objective function given a static dataset of input-output queries. A popular approach to solving this problem is maintaining a proxy model, e.g., a deep neural network (DNN), that approximates the true objective function. Here, the…

Cited by 46SourcePDFScholar
2021

Scaling Neural Tangent Kernels via Sketching and Random Features

NeurIPS 2021poster

The Neural Tangent Kernel (NTK) characterizes the behavior of infinitely-wide neural networks trained under least squares loss by gradient descent. Recent works also report that NTK regression can outperform finitely-wide neural networks trained on small-scale datasets. However, the computational co…

2021

SmoothMix: Training Confidence-calibrated Smoothed Classifiers for Certified Robustness

NeurIPS 2021poster

Randomized smoothing is currently a state-of-the-art method to construct a certifiably robust classifier from neural networks against $\ell_2$-adversarial perturbations. Under the paradigm, the robustness of a classifier is aligned with the prediction confidence, i.e., the higher confidence from a s…

2021

State Entropy Maximization with Random Encoders for Efficient Exploration

ICML 2021spotlight

Recent exploration methods have proven to be a recipe for improving sample-efficiency in deep reinforcement learning (RL). However, efficient exploration in high-dimensional observation spaces still remains a challenge. This paper presents Random Encoders for Efficient Exploration (RE3), an explorat…

2020

Adversarial Neural Pruning with Latent Vulnerability Suppression

ICML 2020poster

Despite the remarkable performance of deep neural networks on various computer vision tasks, they are known to be susceptible to adversarial perturbations, which makes it challenging to deploy them in real-world safety-critical applications. In this paper, we conjecture that the leading cause of adv…

2020

CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances

NeurIPS 2020poster

Novelty detection, i.e., identifying whether a given sample is drawn from outside the training distribution, is essential for reliable machine learning. To this end, there have been many attempts at learning a representation well-suited for novelty detection and designing a score based on such repre…

2020

Consistency Regularization for Certified Robustness of Smoothed Classifiers

NeurIPS 2020poster

A recent technique of randomized smoothing has shown that the worst-case (adversarial) l2-robustness can be transformed into the average-case Gaussian-robustness by "smoothing" a classifier, i.e., by considering the averaged prediction over Gaussian noise. In this paradigm, one should rethink the no…

2020

Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning

ICML 2020poster

Model-based reinforcement learning (RL) enjoys several benefits, such as data-efficiency and planning, by learning a model of the environment’s dynamics. However, learning a global model that can generalize across different dynamics remains a challenge. To tackle this problem, we decompose the task…

2020

Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised Learning

NeurIPS 2020poster

While semi-supervised learning (SSL) has proven to be a promising way for leveraging unlabeled data when labeled data is scarce, the existing SSL algorithms typically assume that training class distributions are balanced. However, these SSL algorithms trained under imbalanced class distributions can…

2020

Few-shot Visual Reasoning with Meta-Analogical Contrastive Learning

NeurIPS 2020poster

While humans can solve a visual puzzle that requires logical reasoning by observing only few samples, it would require training over a large number of samples for state-of-the-art deep reasoning models to obtain similar performance on the same task. In this work, we propose to solve such a few-shot…

Cited by 29SourcePDFScholar
2020

Guiding Deep Molecular Optimization with Genetic Exploration

NeurIPS 2020poster

De novo molecular design attempts to search over the chemical space for molecules with the desired property. Recently, deep learning has gained considerable attention as a promising approach to solve the problem. In this paper, we propose genetic expert-guided learning (GEGL), a simple yet novel fra…

2020

Learning from Failure: De-biasing Classifier from Biased Classifier

NeurIPS 2020poster

Neural networks often learn to make predictions that overly rely on spurious corre- lation existing in the dataset, which causes the model to be biased. While previous work tackles this issue by using explicit labeling on the spuriously correlated attributes or presuming a particular bias type, we i…

2020

Lookahead: A Far-sighted Alternative of Magnitude-based Pruning

ICLR 2020poster

Magnitude-based pruning is one of the simplest methods for pruning neural networks. Despite its simplicity, magnitude-based pruning and its variants demonstrated remarkable performances for pruning modern architectures. Based on the observation that magnitude-based pruning indeed minimizes the Frobe…

Cited by 127SourcecodeScholar
2020

Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning

ICLR 2020poster

Deep reinforcement learning (RL) agents often fail to generalize to unseen environments (yet semantically similar to trained agents), particularly when they are trained on high-dimensional state spaces, such as images. In this paper, we propose a simple technique to improve a generalization ability…

Cited by 246SourcecodeScholar
2020

Regularizing Class-Wise Predictions via Self-Knowledge Distillation

CVPR 2020poster

Deep neural networks with millions of parameters may suffer from poor generalization due to overfitting. To mitigate the issue, we propose a new regularization method that penalizes the predictive distribution between similar samples. In particular, we distill the predictive distribution between dif…

Cited by 379PDFcodeScholar
2020

Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement Learning

NeurIPS 2020poster

Model-based reinforcement learning (RL) has shown great potential in various control tasks in terms of both sample-efficiency and final performance. However, learning a generalizable dynamics model robust to changes in dynamics remains a challenge since the target transition dynamics follow a multi-…

2019

Iterative Bayesian Learning for Crowdsourced Regression

AISTATS 2019poster

Crowdsourcing platforms emerged as popular venues for purchasing human intelligence at low cost for large volume of tasks. As many low-paid workers are prone to give noisy answers, a common practice is to add redundancy by assigning multiple workers to each task and then simply average out these ans…

Cited by 9SourcePDFScholar
2019

Mining GOLD Samples for Conditional GANs

NeurIPS 2019poster

Conditional generative adversarial networks (cGANs) have gained a considerable attention in recent years due to its class-wise controllability and superior quality for complex generation tasks. We introduce a simple yet effective approach to improving cGANs by measuring the discrepancy between the d…

2019

Overcoming Catastrophic Forgetting With Unlabeled Data in the Wild

ICCV 2019poster

Lifelong learning with deep neural networks is well-known to suffer from catastrophic forgetting: the performance on previous tasks drastically degrades when learning a new task. To alleviate this effect, we propose to leverage a large stream of unlabeled data easily obtainable in the wild. In parti…

Cited by 285PDFcodeScholar
2019

Robust Inference via Generative Classifiers for Handling Noisy Labels

ICML 2019oral

Large-scale datasets may contain significant proportions of noisy (incorrect) class labels, and it is well-known that modern deep neural networks (DNNs) poorly generalize from such noisy training datasets. To mitigate the issue, we propose a novel inference method, termed Robust Generative classifie…

2018

A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks

NeurIPS 2018spotlight

Detecting test samples drawn sufficiently far away from the training distribution statistically or adversarially is a fundamental requirement for deploying a good classifier in many real-world machine learning applications. However, deep neural networks with the softmax classifier are known to produ…

2018

Gauged Mini-Bucket Elimination for Approximate Inference

AISTATS 2018poster

Computing the partition function Z of a discrete graphical model is a fundamental inference challenge. Since this is computationally intractable, variational approximations are often used in practice. Recently, so-called gauge transformations were used to improve variational lower bounds on Z. In th…

Cited by 0SourcePDFScholar
2018

Hierarchical Novelty Detection for Visual Object Recognition

CVPR 2018poster

Deep neural networks have achieved impressive success in large-scale visual object recognition tasks with a predefined set of classes. However, recognizing objects of novel classes unseen during training still remains challenging. The problem of detecting such novel classes has been addressed in the…

Cited by 94SourcePDFScholar
2018

Learning to Specialize with Knowledge Distillation for Visual Question Answering

NeurIPS 2018poster

Visual Question Answering (VQA) is a notoriously challenging problem because it involves various heterogeneous tasks defined by questions within a unified framework. Learning specialized models for individual types of tasks is intuitively attracting but surprisingly difficult; it is not straightforw…

2018

Training Confidence-calibrated Classifiers for Detecting Out-of-Distribution Samples

ICLR 2018poster

The problem of detecting whether a test sample is from in-distribution (i.e., training distribution by a classifier) or out-of-distribution sufficiently different from it arises in many real-world machine learning applications. However, the state-of-art deep neural networks are known to be highly ov…

2017

Faster Greedy MAP Inference for Determinantal Point Processes

ICML 2017poster

Determinantal point processes (DPPs) are popular probabilistic models that arise in many machine learning tasks, where distributions of diverse sets are characterized by determinants of their features. In this paper, we develop fast algorithms to find the most likely configuration (MAP) of large-sca…

2017

Rapid Mixing Swendsen-Wang Sampler for Stochastic Partitioned Attractive Models

AISTATS 2017poster

The Gibbs sampler is the most popular Markov chain used for learning and inference problems in Graphical Models (GM). These tasks are computationally intractable in general, and the Gibbs sampler often suffers from slow mixing. In this paper, we study the Swendsen-Wang dynamics which is a more sop…

Cited by 9SourcePDFScholar
2015

Large-scale log-determinant computation through stochastic Chebyshev expansions

ICML 2015poster

Logarithms of determinants of large positive definite matrices appear ubiquitously in machine learning applications including Gaussian graphical and Gaussian process models, partition functions of discrete graphical models, minimum-volume ellipsoids and metric and kernel learning. Log-determinant co…

Cited by 121SourcePDFScholar
2015

Minimum Weight Perfect Matching via Blossom Belief Propagation

NeurIPS 2015spotlight

Max-product Belief Propagation (BP) is a popular message-passing algorithm for computing a Maximum-A-Posteriori (MAP) assignment over a distribution represented by a Graphical Model (GM). It has been shown that BP can solve a number of combinatorial optimization problems including minimum weight mat…

Cited by 9SourcePDFScholar