← Search

Jaegul Choo

112 accepted papers

2026

ACG: Action Coherence Guidance for Flow-Based Vision-Language-Action Models

ICRA 2026poster

Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet, when trained via imitation learning, their high generative capacity makes them sensitive to noise in human demonstration…

2026

AHS: Adaptive Head Synthesis via Synthetic Data Augmentations

CVPR 2026

Recent digital media advancements have created increasing demands for sophisticated portrait manipulation techniques, particularly head swapping--where one image's head is seamlessly integrated onto another's body. Current approaches predominantly rely on face-centered cropped data with limited view

Cited by 0SourcecodeScholar
2026

Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

IJCAI 2026

Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware cac

Cited by 0Scholar
2026

Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation

CVPR 2026

This paper addresses the challenge of data scarcity in semantic segmentation by generating datasets through text-to-image (T2I) generation models, reducing image acquisition and labeling costs. Segmentation dataset generation faces two key challenges: 1) aligning generated samples with the target do

Cited by 0SourcecodeScholar
2026

EgoX: Egocentric Video Generation from a Single Exocentric Video

CVPR 2026

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive understanding but remains highly challenging due to extreme c

Cited by 0SourcecodeScholar
2026

ExpGuard: LLM Content Moderation in Specialized Domains

ICLR 2026poster

With the growing deployment of large language models (LLMs) in real-world applications, establishing robust safety guardrails to moderate their inputs and outputs has become essential to ensure adherence to safety policies. Current guardrail models predominantly address general human-LLM interaction…

Cited by 0SourcecodeScholar
2026

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control

RSS 2026poster

Simulation-based reinforcement learning (RL) is central for robotic control when expert demonstrations are unavailable. However, scaling RL to high-dimensional robots remains challenging. On-policy methods such as PPO are reliable but require large amounts of simulation because they discard past dat…

Cited by 0SourceScholar
2026

LiveWeb-IE: A Benchmark For Online Web Information Extraction

ICLR 2026poster

Web information extraction (WIE) is the task of automatically extracting data from web pages, offering high utility for various applications. The evaluation of WIE systems has traditionally relied on benchmarks built from HTML snapshots captured at a single point in time. However, this offline evalu…

Cited by 0SourceScholar
2026

MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning

CVPR 2026

Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal language models (MLLMs) for this purpose, but their substantial computational cost hinders practical application.This limi

Cited by 0SourcecodeScholar
2026

Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping

CVPR 2026

Diffusion Transformers (DiTs) have significantly enhanced text-to-image (T2I) generation quality, enabling high-quality personalized content creation. However, fine-tuning these models requires substantial computational complexity and memory, limiting practical deployment under resource constraints.

Cited by 0SourceScholar
2026

More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences

ICLR 2026poster

Recommender systems traditionally rely on the principle of Revealed Preference (RP), which assumes that observed user behaviors faithfully reflect underlying interests. While effective at scale, this assumption is fragile in practice, as real-world choices are often noisy and inconsistent. Thus, eve…

Cited by 0SourcecodeScholar
2026

OPRO: Orthogonal Panel-Relative Operators for Panel-Aware In-Context Image Generation

CVPR 2026

We introduce a parameter-efficient adaptation method for panel-aware in-context image generation with pre-trained diffusion transformers. The key idea is to compose learnable, panel-specific orthogonal operators onto the backbone's frozen positional encodings. This design provides two desirable prop

Cited by 0SourceScholar
2026

Position: Significant impact of numerical precision in scientific machine learning

ICML 2026poster

The machine learning community has focused on computational efficiency, often leveraging lower-precision formats such as FP16, rather than the standard FP32. In contrast, little attention has been paid to higher-precision formats, such as FP64, despite their critical role in scientific domains like …

Cited by 0SourceScholar
2026

Selectively Extracting and Injecting Visual Attributes into Text-to-Image Models

CVPR 2026

Text-to-image models are increasingly utilized in design workflows, but articulating nuanced design intentions solely through text remains a challenge. This work proposes a method that extracts visual attributes from a reference image and injects them directly into the generation pipeline. Specifica

Cited by 0SourceScholar
2026

Sparsity-promoting Fine-tuning for Equivariant Materials Foundation Model

ICLR 2026poster

Pre-trained materials foundation models, or machine learning interatomic potentials, leverage general physicochemical knowledge to effectively approximate potential energy surfaces. However, they often require domain-specific calibration due to physicochemical diversity and mismatches between practi…

Cited by 0SourceScholar
2026

SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation

AAAI 2026technical

The increasing demand for AR/VR applications has highlighted the need for high-quality content, such as 360° live wallpapers. However, generating high-quality 360° panoramic contents remains a challenging task due to the severe distortions introduced by equirectangular projection (ERP). Existing ap

Cited by 0SourcePDFScholar
2026

Think Wise, Collaborate Effectively: A Rationale-Aware LLM-Based Recommender with Reinforcement Learning from Collaborative Signals

AAAI 2026technical

Large Language Models (LLMs) have recently emerged as powerful reasoning engines in recommender systems, generating natural-language explanations that foster user engagement. However, their recommendation performance remains limited, as they lack exposure to collaborative user-item interaction patte

Cited by 0SourcePDFScholar
2025

ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts

NeurIPS 2025poster

Dataset bias, where data points are skewed to certain concepts, is ubiquitous in machine learning datasets. Yet, systematically identifying these biases is challenging without costly, fine-grained attribute annotations. We present ConceptScope, a scalable and automated framework for analyzing visual…

Cited by 0SourceScholar
2025

Delving into Large Language Models for Effective Time-Series Anomaly Detection

NeurIPS 2025poster

Recent efforts to apply Large Language Models (LLMs) to time-series anomaly detection (TSAD) have yielded limited success, often performing worse than even simple methods. While prior work has focused solely on downstream performance evaluation, the fundamental question—why do LLMs struggle with TSA…

Cited by 0SourcecodeScholar
2025

Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention

CVPR 2025poster

While large-scale text-to-image diffusion models enable the generation of high-quality, diverse images from text prompts, these prompts struggle to capture intricate details, such as textures, preventing the user intent from being reflected. This limitation has led to efforts to generate images cond…

Cited by 0SourcePDFScholar
2025

Enabling Region-Specific Control via Lassos in Point-Based Colorization

AAAI 2025technical

Point-based interactive colorization techniques allow users to effortlessly colorize grayscale images using user-provided color hints. However, point-based methods often face challenges when different colors are given to semantically similar areas, leading to color intermingling and unsatisfactory r…

Cited by 0SourcePDFScholar
2025

Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration

ACL 2025long

To create culturally inclusive vision-language models (VLMs), developing a benchmark that tests their ability to address culturally relevant questions is essential. Existing approaches typically rely on human annotators, making the process labor-intensive and creating a cognitive burden in generatin…

Cited by 0SourcePDFScholar
2025

Exploring In-context Example Generation for Machine Translation

ACL 2025finding

Large language models (LLMs) have demonstrated strong performance across various tasks, leveraging their exceptional in-context learning ability with only a few examples.Accordingly, the selection of optimal in-context examples has been actively studied in the field of machine translation.However, t…

2025

Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free Attention

ICCV 2025poster

Recent advancements in diffusion-based text-to-image (T2I) models have enabled the generation of high-quality and photorealistic images from text. However, they often exhibit societal biases related to gender, race, and socioeconomic status, thereby potentially reinforcing harmful stereotypes and sh…

Cited by 0SourcePDFScholar
2025

Hyperspherical Normalization for Scalable Deep Reinforcement Learning

ICML 2025spotlight

Scaling up the model size and computation has brought consistent performance improvements in supervised learning. However, this lesson often fails to apply to reinforcement learning (RL) because training the model on non-stationary data easily leads to overfitting and unstable optimization. In resp…

Cited by 0SourcePDFScholar
2025

Opt-Out: Investigating Entity-Level Unlearning for Large Language Models via Optimal Transport

ACL 2025long

Instruction-following large language models (LLMs), such as ChatGPT, have become widely popular among everyday users. However, these models inadvertently disclose private, sensitive information to their users, underscoring the need for machine unlearning techniques to remove selective information fr…

Cited by 0SourcePDFScholar
2025

PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask

ICCV 2025poster

Recent virtual try-on approaches have advanced by finetuning pre-trained text-to-image diffusion models to leverage their powerful generative ability; however, the use of text prompts in virtual try-on remains underexplored. This paper tackles a text-editable virtual try-on task that modifies the cl…

2025

Revisiting LLMs as Zero-Shot Time Series Forecasters: Small Noise Can Break Large Models

ACL 2025short

Large Language Models (LLMs) have shown remarkable performance across diverse tasks without domain-specific training, fueling interest in their potential for time-series forecasting. While LLMs have shown potential in zero-shot forecasting through prompting alone, recent studies suggest that LLMs la…

2025

Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs

EMNLP 2025

Masked diffusion models (MDMs) offer a promising non-autoregressive alternative for large language modeling. Standard decoding methods for MDMs, such as confidence-based sampling, select tokens independently based on individual token confidences at each diffusion step. However, we observe that this

Cited by 0SourcePDFScholar
2025

SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

ICLR 2025spotlight

Recent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to overfitting. These large networks avoid overfitting by integrating components that induce a simplicity bias, guiding models…

2025

Single Ground Truth Is Not Enough: Adding Flexibility to Aspect-Based Sentiment Analysis Evaluation

NAACL 2025long

Aspect-based sentiment analysis (ABSA) is a challenging task of extracting sentiments along with their corresponding aspects and opinion terms from the text.The inherent subjectivity of span annotation makes variability in the surface forms of extracted terms, complicating the evaluation process.Tra…

2025

Sparse autoencoders reveal selective remapping of visual concepts during adaptation

ICLR 2025poster

Adapting foundation models for specific purposes has become a standard approach to build machine learning systems for downstream applications. Yet, it is an open question which mechanisms take place during adaptation. Here we develop a new Sparse Autoencoder (SAE) for the CLIP vision transformer, na…

2025

Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling

CVPR 2025poster

Diffusion models have emerged as a powerful tool for generating high-quality images, videos, and 3D content. While sampling guidance techniques like CFG improve quality, they reduce diversity and motion. Autoguidance mitigates these issues but demands extra weak model training, limiting its practica…

2025

SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head Avatars

ICLR 2025poster

Recent advancements in head avatar rendering using Gaussian primitives have achieved significantly high-fidelity results. Although precise head geometry is crucial for applications like mesh reconstruction and relighting, current methods struggle to capture intricate geometric details and render uns…

Cited by 1SourcePDFScholar
2025

Temporal In‑Context Fine‑Tuning for Versatile Control of Video Diffusion Models

NeurIPS 2025poster

Recent advances in text-to-video diffusion models have enabled high-quality video synthesis, but controllable generation remains challenging—particularly under limited data and compute. Existing fine-tuning methods often rely on external encoders or architectural modifications, which demand large da…

Cited by 0SourceScholar
2025

What to Preserve and What to Transfer: Faithful, Identity-Preserving Diffusion-based Hairstyle Transfer

AAAI 2025technical

Hairstyle transfer is a challenging task in the image editing field that modifies the hairstyle of a given face image while preserving its other appearance and background features. The existing hairstyle transfer approaches heavily rely on StyleGAN, which is pre-trained on cropped and aligned face i…

2024

Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control

ICML 2024poster

Vision Transformers (ViT), when paired with large-scale pretraining, have shown remarkable performance across various computer vision tasks, primarily due to their weak inductive bias. However, while such weak inductive bias aids in pretraining scalability, this may hinder the effective adaptation o…

2024

Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models

EMNLP 2024finding

Pretrained language models memorize vast amounts of information, including private and copyrighted data, raising significant safety concerns. Retraining these models after excluding sensitive data is prohibitively expensive, making machine unlearning a viable, cost-effective alternative. Previous re…

2024

Do's and Don'ts: Learning Desirable Skills with Instruction Videos

NeurIPS 2024poster

Unsupervised skill discovery is a learning paradigm that aims to acquire diverse behaviors without explicit rewards. However, it faces challenges in learning complex behaviors and often leads to learning unsafe or undesirable behaviors. For instance, in various continuous control tasks, current unsu…

2024

EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have demonstrated remarkable in-context learning capabilities across diverse applications. In this work, we explore the effectiveness of LLMs for generating realistic synthetic tabular data, identifying key prompt design elements to optimize performance. We introduce EPI…

2024

Effective Rank Analysis and Regularization for Enhanced 3D Gaussian Splatting

NeurIPS 2024poster

3D reconstruction from multi-view images is one of the fundamental challenges in computer vision and graphics. Recently, 3D Gaussian Splatting (3DGS) has emerged as a promising technique capable of real-time rendering with high-quality 3D reconstruction. This method utilizes 3D Gaussian representati…

2024

Enhancing Intrinsic Features for Debiasing via Investigating Class-Discerning Common Attributes in Bias-Contrastive Pair

CVPR 2024poster

In the image classification task deep neural networks frequently rely on bias attributes that are spuriously correlated with a target class in the presence of dataset bias resulting in degraded performance when applied to data without bias attributes. The task of debiasing aims to compel classifiers…

Cited by 0SourcePDFScholar
2024

Expression Domain Translation Network for Cross-Domain Head Reenactment

ICASSP 2024accepted

Despite the remarkable advancements in head reenactment, the existing methods face challenges in cross-domain head reenactment, which aims to transfer human motions to domains outside the human, including cartoon characters. It is still difficult to extract motion from out-of-domain images due to th…

Cited by 0SourceScholar
2024

Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling

EMNLP 2024finding

Predicting future international events from textual information, such as news articles, has tremendous potential for applications in global policy, strategic decision-making, and geopolitics. However, existing datasets available for this task are often limited in quality, hindering the progress of r…

2024

Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning

ICML 2024poster

Recently, various pre-training methods have been introduced in vision-based Reinforcement Learning (RL). However, their generalization ability remains unclear due to evaluations being limited to in-distribution environments and non-unified experimental setups. To address this, we introduce the Atari…

2024

Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization

ECCV 2024poster

"The task of personalized image aesthetic assessment seeks to tailor aesthetic score prediction models to match individual preferences with just a few user-provided inputs. However, the scalability and generalization capabilities of current approaches are considerably restricted by their reliance on…

2024

Self-Supervised Contrastive Learning for Long-term Forecasting

ICLR 2024poster

Long-term forecasting presents unique challenges due to the time and memory complexity of handling long sequences. Existing methods, which rely on sliding windows to process long sequences, struggle to effectively capture long-term variations that are partially caught within the short window (i.e.,…

2024

Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks

ICML 2024poster

This study investigates the loss of generalization ability in neural networks, revisiting warm-starting experiments from Ash & Adams. Our empirical analysis reveals that common methods designed to enhance plasticity by maintaining trainability provide limited benefits to generalization. While reinit…

2024

Towards Calibrated Robust Fine-Tuning of Vision-Language Models

NeurIPS 2024poster

Improving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for re…

2024

Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering

ACL 2024findings

Building a reliable visual question answering (VQA) system across different languages is a challenging problem, primarily due to the lack of abundant samples for training. To address this challenge, recent studies have employed machine translation systems for the cross-lingual VQA task. This involve…

2024

When Model Meets New Normals: Test-Time Adaptation for Unsupervised Time-Series Anomaly Detection

AAAI 2024technical

Time-series anomaly detection deals with the problem of detecting anomalous timesteps by learning normality from the sequence of observations. However, the concept of normality evolves over time, leading to a "new normal problem", where the distribution of normality can be changed due to the distrib…

2024

iDet3D: Towards Efficient Interactive Object Detection for LiDAR Point Clouds

AAAI 2024technical

Accurately annotating multiple 3D objects in LiDAR scenes is laborious and challenging. While a few previous studies have attempted to leverage semi-automatic methods for cost-effective bounding box annotation, such methods have limitations in efficiently handling numerous multi-class objects. To ef…

Cited by 5SourcePDFScholar
2023

3D-Aware Generative Model for Improved Side-View Image Synthesis

ICCV 2023poster

While recent 3D-aware generative models have shown photo-realistic image synthesis with multi-view consistency, the synthesized image quality degrades depending on the camera pose (e.g., a face with a blurry and noisy boundary at a side viewpoint). Such degradation is mainly caused by the difficulty…

Cited by 5PDFScholar
2023

AniEE: A Dataset of Animal Experimental Literature for Event Extraction

EMNLP 2023long findings

Event extraction (EE), as a crucial information extraction (IE) task, aims to identify event triggers and their associated arguments from unstructured text, subsequently classifying them into pre-defined types and roles. In the biomedical domain, EE is widely used to extract complex structures repre…

Cited by 0SourceScholar
2023

CAFA: Class-Aware Feature Alignment for Test-Time Adaptation

ICCV 2023poster

Despite recent advancements in deep learning, deep neural networks continue to suffer from performance degradation when applied to new data that differs from training data. Test-time adaptation (TTA) aims to address this challenge by adapting a model to unlabeled data at test time. TTA can be applie…

Cited by 25PDFScholar
2023

DEnsity: Open-domain Dialogue Evaluation Metric using Density Estimation

ACL 2023findings

Despite the recent advances in open-domain dialogue systems, building a reliable evaluation metric is still a challenging problem. Recent studies proposed learnable metrics based on classification models trained to distinguish the correct response. However, neural classifiers are known to make overl…

2023

FaceCLIPNeRF: Text-driven 3D Face Manipulation using Deformable Neural Radiance Fields

ICCV 2023poster

As recent advances in Neural Radiance Fields (NeRF) have enabled high-fidelity 3D face reconstruction and novel view synthesis, its manipulation also became an essential task in 3D vision. However, existing manipulation methods require extensive human labor, such as a user-provided semantic mask and…

Cited by 16PDFcodeScholar
2023

HistRED: A Historical Document-Level Relation Extraction Dataset

ACL 2023long

Despite the extensive applications of relation extraction (RE) tasks in various domains, little has been explored in the historical context, which contains promising data across hundreds and thousands of years. To promote the historical RE research, we present HistRED constructed from Yeonhaengnok.…

2023

Label Shift Adapter for Test-Time Adaptation under Covariate and Label Shifts

ICCV 2023poster

Test-time adaptation (TTA) aims to adapt a pre-trained model to the target domain in a batch-by-batch manner during inference. While label distributions often exhibit imbalances in real-world scenarios, most previous TTA approaches typically assume that both source and target domain datasets have ba…

Cited by 22PDFScholar
2023

Learning to Discover Skills through Guidance

NeurIPS 2023poster

In the field of unsupervised skill discovery (USD), a major challenge is limited exploration, primarily due to substantial penalties when skills deviate from their initial trajectories. To enhance exploration, recent methodologies employ auxiliary rewards to maximize the epistemic uncertainty or ent…

2023

Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image Synthesis

ICCV 2023poster

Existing text-to-image generation approaches have set high standards for photorealism and text-image correspondence, largely benefiting from web-scale text-image datasets, which can include up to 5 billion pairs. However, text-to-image generation models trained on domain-specific datasets, such as u…

Cited by 12PDFcodeScholar
2023

Local 3D Editing via 3D Distillation of CLIP Knowledge

CVPR 2023poster

3D content manipulation is an important computer vision task with many real-world applications (e.g., product design, cartoon generation, and 3D Avatar editing). Recently proposed 3D GANs can generate diverse photo-realistic 3D-aware contents using Neural Radiance fields (NeRF). However, manipulatio…

Cited by 29SourcePDFScholar
2023

On the Importance of Feature Decorrelation for Unsupervised Representation Learning in Reinforcement Learning

ICML 2023poster

Recently, unsupervised representation learning (URL) has improved the sample efficiency of Reinforcement Learning (RL) by pretraining a model from a large unlabeled dataset. The underlying principle of these methods is to learn temporally predictive representations by predicting future states in the…

2023

PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning

NeurIPS 2023poster

In Reinforcement Learning (RL), enhancing sample efficiency is crucial, particularly in scenarios when data acquisition is costly and risky. In principle, off-policy RL algorithms can improve sample efficiency by allowing multiple updates per environment interaction. However, these multiple updates…

2023

Perspective Projection-Based 3d CT Reconstruction from Biplanar X-Rays

ICASSP 2023accepted

X-ray computed tomography (CT) is one of the most common imaging techniques used to diagnose various diseases in the medical field. Its high contrast sensitivity and spatial resolution allow the physician to observe details of body parts such as bones, soft tissue, blood vessels, etc. As it involves…

Cited by 0SourceScholar
2023

Revisiting the Importance of Amplifying Bias for Debiasing

AAAI 2023technical

In image classification, debiasing aims to train a classifier to be less susceptible to dataset bias, the strong correlation between peripheral attributes of data samples and a target class. For example, even if the frog class in the dataset mainly consists of frog images with a swamp background (i.…

Cited by 25SourcePDFScholar
2023

Shortcut-V2V: Compression Framework for Video-to-Video Translation Based on Temporal Redundancy Reduction

ICCV 2023poster

Video-to-video translation aims to generate video frames of a target domain from an input video. Despite its usefulness, the existing networks require enormous computations, necessitating their model compression for wide use. While there exist compression methods that improve computational efficienc…

Cited by 2PDFcodeScholar
2023

SimCKP: Simple Contrastive Learning of Keyphrase Representations

EMNLP 2023long findings

Keyphrase generation (KG) aims to generate a set of summarizing words or phrases given a source document, while keyphrase extraction (KE) aims to identify them from the text. Because the search space is much smaller in KE, it is often combined with KG to predict keyphrases that may or may not exist…

Cited by 0SourcecodeScholar
2023

TTN: A Domain-Shift Aware Batch Normalization in Test-Time Adaptation

ICLR 2023poster

This paper proposes a novel batch normalization strategy for test-time adaptation. Recent test-time adaptation methods heavily rely on the modified batch normalization, i.e., transductive batch normalization (TBN), which calculates the mean and the variance from the current test batch rather than us…

Cited by 111SourcePDFScholar
2023

Towards Accurate Translation via Semantically Appropriate Application of Lexical Constraints

ACL 2023findings

Lexically-constrained NMT (LNMT) aims to incorporate user-provided terminology into translations. Despite its practical advantages, existing work has not evaluated LNMT models under challenging real-world conditions. In this paper, we focus on two important but understudied issues that lie in the cu…

2023

Towards Formality-Aware Neural Machine Translation by Leveraging Context Information

EMNLP 2023short findings

Formality is one of the most important linguistic properties to determine the naturalness of translation. Although a target-side context contains formality-related tokens, the sparsity within the context makes it difficult for context-aware neural machine translation (NMT) models to properly discern…

Cited by 0SourceScholar
2023

Towards Open-Set Test-Time Adaptation Utilizing the Wisdom of Crowds in Entropy Minimization

ICCV 2023poster

Test-time adaptation (TTA) methods, which generally rely on the model's predictions (e.g., entropy minimization) to adapt the source pretrained model to the unlabeled target domain, suffer from noisy signals originating from 1) incorrect or 2) open-set predictions. Long-term stable adaptation is ham…

Cited by 26PDFScholar
2023

Towards Physically Reliable Molecular Representation Learning

UAI 2023poster

Estimating the energetic properties of molecular systems is a critical task in material design. Machine learning has shown remarkable promise on this task over classical force fields, but a fully data-driven approach suffers from limited labeled data; not just the amount of available data lacks, but…

Cited by 2SourcePDFScholar
2022

AnimeCeleb: Large-Scale Animation CelebHeads Dataset for Head Reenactment

ECCV 2022poster

"We present a novel Animation CelebHeads dataset (AnimeCeleb) to address an animation head reenactment. Different from previous animation head datasets, we utilize a 3D animation models as the controllable image samplers, which can provide a large amount of head images with their corresponding detai…

2022

CEDe: A collection of expert-curated datasets with atom-level entity annotations for Optical Chemical Structure Recognition

NeurIPS 2022accept

Optical Chemical Structure Recognition (OCSR) deals with the translation from chemical images to molecular structures, this being the main way chemical compounds are depicted in scientific documents. Traditionally, rule-based methods have followed a framework based on the detection of chemical entit…

Cited by 11SourcePDFScholar
2022

High-Resolution Virtual Try-On with Misalignment and Occlusion-Handled Conditions

ECCV 2022poster

"Image-based virtual try-on aims to synthesize an image of a person wearing a given clothing item. To solve the task, the existing methods warp the clothing item to fit the person’s body and generate the segmentation map of the person wearing the item before fusing the item with the person. However,…

2022

Mining Multi-Label Samples from Single Positive Labels

NeurIPS 2022accept

Conditional generative adversarial networks (cGANs) have shown superior results in class-conditional generation tasks. To simultaneously control multiple conditions, cGANs require multi-label training datasets, where multiple labels can be assigned to each data instance. Nevertheless, the tremendous…

Cited by 7SourcePDFScholar
2022

Pneg: Prompt-based Negative Response Generation for Dialogue Response Selection Task

EMNLP 2022main

In retrieval-based dialogue systems, a response selection model acts as a ranker to select the most appropriate response among several candidates. However, such selection models tend to rely on context-response content similarity, which makes models vulnerable to adversarial responses that are seman…

Cited by 6SourcePDFScholar
2022

Rethinking Style Transformer with Energy-based Interpretation: Adversarial Unsupervised Style Transfer using a Pretrained Model

EMNLP 2022main

Style control, content preservation, and fluency determine the quality of text style transfer models. To train on a nonparallel corpus, several existing approaches aim to deceive the style discriminator with an adversarial loss. However, adversarial training significantly degrades fluency compared t…

Cited by 1SourcePDFScholar
2022

Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift

ICLR 2022poster

Statistical properties such as mean and variance often change over time in time series, i.e., time-series data suffer from a distribution shift problem. This change in temporal distribution is one of the main challenges that prevent accurate time-series forecasting. To address this issue, we propose…

2022

Reweighting Strategy Based on Synthetic Data Identification for Sentence Similarity

COLING 2022main

Semantically meaningful sentence embeddings are important for numerous tasks in natural language processing. To obtain such embeddings, recent studies explored the idea of utilizing synthetically generated data from pretrained language models(PLMs) as a training corpus. However, PLMs often generate…

2022

Style Your Hair: Latent Optimization for Pose-Invariant Hairstyle Transfer via Local-Style-Aware Hair Alignment

ECCV 2022poster

"Editing hairstyle is unique and challenging due to the complexity and delicacy of hairstyle. Although recent approaches significantly improved the hair details, these models often produce undesirable outputs when a pose of a source image is considerably different from that of a target hair image, l…

2022

WaveBound: Dynamic Error Bounds for Stable Time Series Forecasting

NeurIPS 2022accept

Time series forecasting has become a critical task due to its high practicality in real-world applications such as traffic, energy consumption, economics and finance, and disease analysis. Recent deep-learning-based approaches have shown remarkable success in time series forecasting. Nonetheless, du…

Cited by 6SourcePDFScholar
2021

AVocaDo: Strategy for Adapting Vocabulary to Downstream Domain

EMNLP 2021main

During the fine-tuning phase of transfer learning, the pretrained vocabulary remains unchanged, while model parameters are updated. The vocabulary generated based on the pretrained data is suboptimal for downstream data when domain discrepancy exists. We propose to consider the vocabulary as an opti…

2021

Constructing Multi-Modal Dialogue Dataset by Replacing Text with Semantically Relevant Images

ACL 2021short

In multi-modal dialogue systems, it is important to allow the use of images as part of a multi-turn conversation. Training such dialogue systems generally requires a large-scale dataset consisting of multi-turn dialogues that involve images, but such datasets rarely exist. In response, this paper pr…

2021

Deep Edge-Aware Interactive Colorization Against Color-Bleeding Effects

ICCV 2021poster

Deep neural networks for automatic image colorization often suffer from the color-bleeding artifact, a problematic color spreading near the boundaries between adjacent objects. Such color-bleeding artifacts debase the reality of generated outputs, limiting the applicability of colorization models in…

Cited by 41PDFScholar
2021

Efficient Adversarial Audio Synthesis VIA Progressive Upsampling

ICASSP 2021accepted

This paper proposes a novel generative model called PUGAN, which progressively synthesizes high-quality audio in a raw waveform. Progressive upsampling GAN (PUGAN) leverages the progressive generation of higher-resolution output by stacking multiple encoder-decoder architectures. Compared to an exis…

Cited by 0SourceScholar
2021

Learning Debiased Representation via Disentangled Feature Augmentation

NeurIPS 2021oral

Image classification models tend to make decisions based on peripheral attributes of data items that have strong correlation with a target variable (i.e., dataset bias). These biased models suffer from the poor generalization capability when evaluated on unbiased datasets. Existing approaches for de…

Cited by 170SourcePDFScholar
2021

Not Just Compete, but Collaborate: Local Image-to-Image Translation via Cooperative Mask Prediction

CVPR 2021poster

Facial attribute editing aims to manipulate the image with the desired attribute while preserving the other details. Recently, generative adversarial networks along with the encoder-decoder architecture have been utilized for this task owing to their ability to create realistic images. However, the…

Cited by 13PDFcodeScholar
2021

Novel Natural Language Summarization of Program Code via Leveraging Multiple Input Representations

EMNLP 2021finding

The lack of description of a given program code acts as a big hurdle to those developers new to the code base for its understanding. To tackle this problem, previous work on code summarization, the task of automatically generating code description given a piece of code reported that an auxiliary lea…

Cited by 14SourcePDFScholar
2021

Restoring and Mining the Records of the Joseon Dynasty via Neural Language Modeling and Machine Translation

NAACL 2021long

Understanding voluminous historical records provides clues on the past in various aspects, such as social and political issues and even natural science facts. However, it is generally difficult to fully utilize the historical records, since most of the documents are not written in a modern language…

2021

RobustNet: Improving Domain Generalization in Urban-Scene Segmentation via Instance Selective Whitening

CVPR 2021poster

Enhancing the generalization capability of deep neural networks to unseen domains is crucial for safety-critical applications in the real world such as autonomous driving. To address this issue, this paper proposes a novel instance selective whitening loss to improve the robustness of the segmentati…

Cited by 343PDFcodeScholar
2021

Standardized Max Logits: A Simple yet Effective Approach for Identifying Unexpected Road Obstacles in Urban-Scene Segmentation

ICCV 2021poster

Identifying unexpected objects on roads in semantic segmentation (e.g., identifying dogs on roads) is crucial in safety-critical applications. Existing approaches use images of unexpected objects from external datasets or require additional training (e.g., retraining segmentation networks or trainin…

Cited by 113PDFcodeScholar
2021

Unsupervised Neural Machine Translation for Low-Resource Domains via Meta-Learning

ACL 2021long

Unsupervised machine translation, which utilizes unpaired monolingual corpora as training data, has achieved comparable performance against supervised machine translation. However, it still suffers from data-scarce domains. To address this issue, this paper presents a novel meta-learning algorithm f…

2021

VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization

CVPR 2021poster

The task of image-based virtual try-on aims to transfer a target clothing item onto the corresponding region of a person, which is commonly tackled by fitting the item to the desired body part and fusing the warped item with the person. While an increasing number of studies have been conducted, the…

Cited by 302PDFcodeScholar
2021

Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential Equation

AAAI 2021technical

Video generation models often operate under the assumption of fixed frame rates, which leads to suboptimal performance when it comes to handling flexible frame rates (e.g., increasing the frame rate of the more dynamic portion of the video as well as handling missing video frames). To resolve the re…

2020

Cars Can't Fly Up in the Sky: Improving Urban-Scene Segmentation via Height-Driven Attention Networks

CVPR 2020poster

This paper exploits the intrinsic features of urban-scene images and proposes a general add-on module, called height-driven attention networks (HANet), for improving semantic segmentation for urban-scene images. It emphasizes informative features or classes selectively according to the vertical posi…

Cited by 242PDFcodeScholar
2020

Learning De-biased Representations with Biased Representations

ICML 2020poster

Many machine learning algorithms are trained and evaluated by splitting data from a single source into training and test sets. While such focus on in-distribution learning scenarios has led to interesting advancement, it has not been able to tell if models are relying on dataset biases as shortcuts…

2020

NeurQuRI: Neural Question Requirement Inspector for Answerability Prediction in Machine Reading Comprehension

ICLR 2020poster

Real-world question answering systems often retrieve potentially relevant documents to a given question through a keyword search, followed by a machine reading comprehension (MRC) step to find the exact answer from them. In this process, it is essential to properly determine whether an answer to the…

Cited by 34SourceScholar
2020

Reference-Based Sketch Image Colorization Using Augmented-Self Reference and Dense Semantic Correspondence

CVPR 2020poster

This paper tackles the automatic colorization task of a sketch image given an already-colored reference image. Colorizing a sketch image is in high demand in comics, animation, and other content creation applications, but it suffers from information scarcity of a sketch image. To address this, a ref…

Cited by 382PDFScholar
2019

Coloring With Limited Data: Few-Shot Colorization via Memory Augmented Networks

CVPR 2019poster

Despite recent advancements in deep learning-based automatic colorization, they are still limited when it comes to few-shot learning. Existing models require a significant amount of training data. To tackle this issue, we present a novel memory-augmented colorization model MemoPainter that can produ…

Cited by 158PDFScholar
2019

Image-To-Image Translation via Group-Wise Deep Whitening-And-Coloring Transformation

CVPR 2019oral

Recently, unsupervised exemplar-based image-to-image translation, conditioned on a given exemplar without the paired data, has accomplished substantial advancements. In order to transfer the information from an exemplar to an input image, existing methods often use a normalization technique, e.g., a…

Cited by 181PDFcodeScholar
2018

Coloring with Words: Guiding Image Colorization Through Text-based Palette Generation

ECCV 2018poster

This paper proposes a novel approach to generate multiple color palettes that reflect the semantics of input text and then colorize a given grayscale image according to the generated color palette. In contrast to existing approaches, our model can understand rich text, whether it is a single word, a…

2018

StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation

CVPR 2018poster

Recent studies have shown remarkable success in image-to-image translation for two domains. However, existing approaches have limited scalability and robustness in handling more than two domains, since different models should be built independently for every pair of image domains. To address this li…

Cited by 5010SourcePDFScholar