← Search

Sungroh Yoon

83 accepted papers

2026

CANDI: Curated Test-Time Adaptation for Multivariate Time-Series Anomaly Detection Under Distribution Shift

AAAI 2026technical

Multivariate time-series anomaly detection (MTSAD) aims to identify deviations from normality in multivariate time-series and is critical in real-world applications. However, in real-world deployments, distribution shifts are ubiquitous and cause severe performance degradation in pre-trained anomaly

Cited by 0SourcePDFScholar
2026

Contextualized Visual Personalization in Vision-Language Models

ICML 2026poster

Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's specific experiences, as they lack the ability to associate visual inputs with a user’s accumulated visual-textual context. We newly formalize this challeng…

Cited by 0SourceScholar
2026

Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

ICML 2026poster

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects or attributes fail to emerge in the generated image. By performing linear probing on text tokens, we demonstrate that t…

Cited by 0SourceScholar
2026

HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT

CVPR 2026

Visual Geometry Grounded Transformer (VGGT) has advanced 3D vision, yet its global attention layers suffer from quadratic computational costs that hinder scalability. Several sparsification-based acceleration techniques have been proposed to alleviate this issue, but they often suffer from substanti

Cited by 0SourcecodeScholar
2026

MobileKGQA: On-Device KGQA System on Dynamic Mobile Environments

ICLR 2026poster

Developing a mobile system capable of generating responses based on stored user data is a crucial challenge. Since user data is stored in the form of Knowledge Graphs, the field of knowledge graph question answering (KGQA) presents a promising avenue towards addressing this problem. However, existin…

Cited by 0SourceScholar
2026

R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning

ICML 2026poster

For reinforcement learning in data-scarce domains like real-world robotics, intensive data reuse enhances efficiency but induces overfitting. While prior works focus on critic bias, representation-level instability in Self-Predictive Learning (SPL) under high Update-to-Data (UTD) regimes remains und…

Cited by 0SourceScholar
2026

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

ICLR 2026poster

Recent advances in sequence modeling have highlighted Mamba as a state space architecture offering efficient long-range dependency modeling and providing a viable alternative to Transformers. Building upon this, Mamba-2 introduces the Structured State Space Duality (SSD), which integrates recurrent…

Cited by 0SourcecodeScholar
2026

Time-Conditioned Foreseeing: An EHR-Specific Foundation Model for Irregular Dynamics and Calendrical Time

ICML 2026poster

Electronic Health Records (EHRs) possess unique characteristics distinct from natural language, yet existing EHR foundation models often rely on suboptimal NLP-based approaches. We propose a pretraining method tailored to EHRs' distinct features. First, we introduce Pathology-Focused Binning, a dens…

Cited by 0SourceScholar
2025

Battling the Non-stationarity in Time Series Forecasting via Test-time Adaptation

AAAI 2025technical

Deep Neural Networks have spearheaded remarkable advancements in time series forecasting (TSF), one of the major tasks in time series modeling. Nonetheless, the non-stationarity of time series undermines the reliability of pre-trained source time series forecasters in mission-critical deployment set…

2025

Causality-Aware Contrastive Learning for Robust Multivariate Time-Series Anomaly Detection

ICML 2025poster

Utilizing the complex inter-variable causal relationships within multivariate time-series provides a promising avenue toward more robust and reliable multivariate time-series anomaly detection (MTSAD) but remains an underexplored area of research. This paper proposes Causality-Aware contrastive lear…

2025

Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment

NAACL 2025long

A binary decision task, like yes-no questions or answer verification, reflects a significant real-world scenario such as where users look for confirmation about the correctness of their decisions on specific issues. In this work, we observe that language models exhibit a negative bias in the binary…

2025

DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection

CVPR 2025highlight

Developing effective visual inspection models remains challenging due to the scarcity of defect data. While image generation models have been used to synthesize defect images, producing highly realistic defects remains difficult. We propose DefectFill, a novel method for realistic defect generation…

Cited by 0SourcePDFScholar
2025

Disentangled Motion Modeling for Video Frame Interpolation

AAAI 2025technical

Video Frame Interpolation (VFI) aims to synthesize intermediate frames between existing frames to enhance visual smoothness and quality. Beyond the conventional methods based on the reconstruction loss, recent works have employed generative models for improved perceptual quality. However, they requi…

2025

Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models

ACL 2025finding

Recent advancements in multi-turn voice interaction models have improved user-model communication. However, while closed-source models effectively retain and recall past utterances, whether open-source models share this ability remains unexplored. To fill this gap, we systematically evaluate how wel…

Cited by 0SourcePDFScholar
2025

EdiText: Controllable Coarse-to-Fine Text Editing with Diffusion Language Models

ACL 2025long

We propose EdiText, a controllable text editing method that modifies the reference text to desired attributes at various scales. We integrate an SDEdit-based editing technique that allows for broad adjustments in the degree of text editing. Additionally, we introduce a novel fine-level editing metho…

2025

Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis

ACL 2025long

Personalized AI assistants, a hallmark of the human-like capabilities of Large Language Models (LLMs), are a challenging application that intertwines multiple problems in LLM research. Despite the growing interest in the development of personalized assistants, the lack of an open-source conversation…

2025

Interpretable Next-token Prediction via the Generalized Induction Head

NeurIPS 2025poster

While large transformer models excel in predictive performance, their lack of interpretability restricts their usefulness in high-stakes domains. To remedy this, we propose the Generalized Induction-Head Model (GIM), an interpretable model for next-token prediction inspired by the observation of “in…

Cited by 0SourcecodeScholar
2025

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

ICCV 2025poster

While CLIP has significantly advanced multimodal understanding by bridging vision and language, the inability to grasp negation -- such as failing to differentiate concepts like "parking" from "no parking" -- poses substantial challenges.By analyzing the data used in the public CLIP model's pre-trai…

Cited by 0SourcePDFScholar
2025

Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator

CVPR 2025poster

Subject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and the semantic content of a text prompt. Traditional methods rely on time- and resource-intensive fine-tuning for subject…

Cited by 13SourcePDFScholar
2025

NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers

ICASSP 2025accepted

We present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker adaptation technique capable of fine-tuning multiple references in parallel, significantly reducing training time. Beyo…

Cited by 0SourceScholar
2025

RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

NeurIPS 2025poster

Recent multi-modal large language models (MLLMs) often struggle to generate personalized image captions, even when trained on high-quality captions. In this work, we observe that such limitations persist in existing post-training-based MLLM personalization methods. Specifically, despite being post-t…

Cited by 0SourcecodeScholar
2025

Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion

CVPR 2025poster

Recent advancements in text-to-image models, such as Stable Diffusion, show significant demographic biases. Existing de-biasing techniques rely heavily on additional training, which imposes high computational costs and risks of compromising core image generation functionality. This hinders them from…

Cited by 3SourcePDFScholar
2025

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

ICML 2025poster

Multimodal large language models (MLLMs) excel at generating highly detailed captions but often produce hallucinations. Our analysis reveals that existing hallucination detection methods struggle with detailed captions. We attribute this to the increasing reliance of MLLMs on their generated text, r…

Cited by 1SourcePDFScholar
2025

Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context

NAACL 2025findings

Multi-hop reasoning, which requires multi-step reasoning based on the supporting documents within a given context, remains challenging for large language models (LLMs). LLMs often struggle to filter out irrelevant documents within the context, and their performance is sensitive to the absolute posit…

Cited by 0SourcePDFScholar
2025

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

ICML 2025poster

Detailed image captioning is essential for tasks like data generation and aiding visually impaired individuals. High-quality captions require a balance between precision and recall, which remains challenging for current multimodal large language models (MLLMs). In this work, we hypothesize that this…

Cited by 0SourcePDFScholar
2025

VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance

ICASSP 2025accepted

When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech sy…

Cited by 0SourceScholar
2024

Controlled Text Generation for Black-box Language Models via Score-based Progressive Editor

ACL 2024long

Controlled text generation, aiming to ensure that language models produce text containing only the desired domain or corpus attributes, is immensely crucial in the practical application of language models. Existing methods, however, are inapplicable to black-box models or suffer a significant trade-…

2024

DAFA: Distance-Aware Fair Adversarial Training

ICLR 2024poster

The disparity in accuracy between classes in standard training is amplified during adversarial training, a phenomenon termed the robust fairness problem. Existing methodologies aimed to enhance robust fairness by sacrificing the model's performance on easier classes in order to improve its performan…

2024

Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors

ICLR 2024spotlight

Test-time adaptation (TTA) fine-tunes pre-trained deep neural networks for unseen test data. The primary challenge of TTA is limited access to the entire test dataset during online updates, causing error accumulation. To mitigate it, TTA methods have utilized the model output's entropy as a confiden…

2024

Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach

ACL 2024long

In this paper, we primarily address the issue of dialogue-form context query within the interactive text-to-image retrieval task. Our methodology, PlugIR, actively utilizes the general instruction-following capability of LLMs in two ways. First, by reformulating the dialogue-form context, we elimina…

2024

Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting

NeurIPS 2024poster

Sequence modeling faces challenges in capturing long-range dependencies across diverse tasks. Recent linear and transformer-based forecasters have shown superior performance in time series forecasting. However, they are constrained by their inherent inability to effectively address long-range depend…

2024

LLM-based Frameworks for API Argument Filling in Task-Oriented Conversational Systems

NAACL 2024industry

Task-orientated conversational agents interact with users and assist them via leveraging external APIs. A typical task-oriented conversational system can be broken down into three phases: external API selection, argument filling, and response generation. The focus of our work is the task of argument…

Cited by 4SourcePDFScholar
2024

Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

NeurIPS 2024poster

Recent work shows promising results in expanding the capabilities of large language models (LLM) to directly understand and synthesize speech. However, an LLM-based strategy for modeling spoken dialogs remains elusive, calling for further investigation. This paper introduces an extensive speech-text…

2024

SF(DA)$^2$: Source-free Domain Adaptation Through the Lens of Data Augmentation

ICLR 2024poster

In the face of the deep learning model's vulnerability to domain shift, source-free domain adaptation (SFDA) methods have been proposed to adapt models to new, unseen target domains without requiring access to source domain data. Although the potential benefits of applying data augmentation to SFDA…

2024

Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP

EMNLP 2024finding

A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through natural language. Despite the varying significance of different textual…

Cited by 0SourcePDFScholar
2024

Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection

NeurIPS 2024poster

In our study, we explore methods for detecting unwanted content lurking in visual datasets. We provide a theoretical analysis demonstrating that a model capable of successfully partitioning visual data can be obtained using only textual data. Based on the analysis, we propose Hassle-Free Textual Tra…

2024

Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization

NeurIPS 2024poster

Estimating the homography between two images is crucial for mid- or high-level vision tasks, such as image stitching and fusion. However, using supervised learning methods is often challenging or costly due to the difficulty of collecting ground-truth data. In response, unsupervised learning approac…

2023

BigVGAN: A Universal Neural Vocoder with Large-Scale Training

ICLR 2023poster

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across various recording environments. In this work, we present BigVGAN,…

2023

CLeAR: Continual Learning on Algorithmic Reasoning for Human-like Intelligence

NeurIPS 2023poster

Continual learning (CL) aims to incrementally learn multiple tasks that are presented sequentially. The significance of CL lies not only in the practical importance but also in studying the learning mechanisms of humans who are excellent continual learners. While most research on CL has been done on…

2023

Improving Visual Prompt Tuning for Self-supervised Vision Transformers

ICML 2023poster

Visual Prompt Tuning (VPT) is an effective tuning method for adapting pretrained Vision Transformers (ViTs) to downstream tasks. It leverages extra learnable tokens, known as prompts, which steer the frozen pretrained ViTs. Although VPT has demonstrated its applicability with supervised vision trans…

2023

Large-scale Lifelong Learning of In-context Instructions and How to Tackle It

ACL 2023long

Jointly fine-tuning a Pre-trained Language Model (PLM) on a pre-defined set of tasks with in-context instructions has been proven to improve its generalization performance, allowing us to build a universal language model that can be deployed across task boundaries. In this work, we explore for the f…

Cited by 15SourcePDFScholar
2023

Model Intrinsic Features of Fine-tuning based Text Summarization Models for Factual Consistency

ACL 2023findings

In this study, we analyze the model intrinsic features of a summarization model by varying the fine-tuning objectives and datasets. We fine-tune BART models combining three fine-tuning objectives (negative log-likelihood, unlikelihood, and contrastive loss) and two datasets (CNN/DailyMail and XSum)…

2023

New Insights for the Stability-Plasticity Dilemma in Online Continual Learning

ICLR 2023poster

The aim of continual learning is to learn new tasks continuously (i.e., plasticity) without forgetting previously learned knowledge from old tasks (i.e., stability). In the scenario of online continual learning, wherein data comes strictly in a streaming manner, the plasticity of online continual le…

2023

On the Impact of Knowledge Distillation for Model Interpretability

ICML 2023poster

Several recent studies have elucidated why knowledge distillation (KD) improves model performance. However, few have researched the other advantages of KD in addition to its improving model performance. In this study, we have attempted to show that KD enhances the interpretability as well as the acc…

2023

On the Powerfulness of Textual Outlier Exposure for Visual OoD Detection

NeurIPS 2023poster

Successful detection of Out-of-Distribution (OoD) data is becoming increasingly important to ensure safe deployment of neural networks. One of the main challenges in OoD detection is that neural networks output overconfident predictions on OoD data, make it difficult to determine OoD-ness of data so…

Cited by 14SourcePDFScholar
2023

P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting

NeurIPS 2023poster

While recent large-scale neural codec language models have shown significant improvement in zero-shot TTS by training on thousands of hours of data, they suffer from drawbacks such as a lack of robustness, slow sampling speed similar to previous autoregressive TTS methods, and reliance on pre-traine…

Cited by 42SourcePDFScholar
2023

PUCA: Patch-Unshuffle and Channel Attention for Enhanced Self-Supervised Image Denoising

NeurIPS 2023poster

Although supervised image denoising networks have shown remarkable performance on synthesized noisy images, they often fail in practice due to the difference between real and synthesized noise. Since clean-noisy image pairs from the real world are extremely costly to gather, self-supervised learning…

Cited by 17SourcePDFScholar
2023

ProPILE: Probing Privacy Leakage in Large Language Models

NeurIPS 2023spotlight

The rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of web-collected data, which may inadvertently include sensitive p…

Cited by 174SourcePDFScholar
2022

AutoSNN: Towards Energy-Efficient Spiking Neural Networks

ICML 2022spotlight

Spiking neural networks (SNNs) that mimic information transmission in the brain can energy-efficiently process spatio-temporal information through discrete and sparse spikes, thereby receiving considerable attention. To improve accuracy and energy efficiency of SNNs, most previous studies have focus…

2022

BNUDC: A Two-Branched Deep Neural Network for Restoring Images From Under-Display Cameras

CVPR 2022poster

The images captured by under-display cameras (UDCs) are degraded by the screen in front of them. We model this degradation in terms of a) diffraction by the pixel grid, which attenuates high-spatial-frequency components of the image; and b) diffuse intensity and color changes caused by the multiple…

Cited by 29PDFScholar
2022

Bridging the Gap Between Classification and Localization for Weakly Supervised Object Localization

CVPR 2022poster

Weakly supervised object localization aims to find a target object region in a given image with only weak supervision, such as image-level labels. Most existing methods use a class activation map (CAM) to generate a localization map; however, a CAM identifies only the most discriminative parts of a…

Cited by 56PDFcodeScholar
2022

Confidence Score for Source-Free Unsupervised Domain Adaptation

ICML 2022spotlight

Source-free unsupervised domain adaptation (SFUDA) aims to obtain high performance in the unlabeled target domain using the pre-trained source model, not the source data. Existing SFUDA methods assign the same importance to all target samples, which is vulnerable to incorrect pseudo-labels. To diffe…

2022

Dataset Condensation with Contrastive Signals

ICML 2022spotlight

Recent studies have demonstrated that gradient matching-based dataset synthesis, or dataset condensation (DC), methods can achieve state-of-theart performance when applied to data-efficient learning tasks. However, in this study, we prove that the existing DC methods can perform worse than the rando…

2022

Demystifying the Neural Tangent Kernel From a Practical Perspective: Can It Be Trusted for Neural Architecture Search Without Training?

CVPR 2022poster

In Neural Architecture Search (NAS), reducing the cost of architecture evaluation remains one of the most crucial challenges. Among a plethora of efforts to bypass training of each candidate architecture to convergence for evaluation, the Neural Tangent Kernel (NTK) is emerging as a promising theore…

Cited by 21PDFcodeScholar
2022

Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance

ICML 2022spotlight

We propose Guided-TTS, a high-quality text-to-speech (TTS) model that does not require any transcript of target speaker using classifier guidance. Guided-TTS combines an unconditional diffusion probabilistic model with a separately trained phoneme classifier for classifier guidance. Our unconditiona…

Cited by 112SourcePDFScholar
2022

Perception Prioritized Training of Diffusion Models

CVPR 2022poster

Diffusion models learn to restore noisy data, which is corrupted with different levels of noise, by optimizing the weighted sum of the corresponding loss terms, i.e., denoising score matching loss. In this paper, we show that restoring data corrupted with certain noise levels offers a proper pretext…

Cited by 254PDFcodeScholar
2022

PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior

ICLR 2022poster

Denoising diffusion probabilistic models have been recently proposed to generate high-quality samples by estimating the gradient of the data density. The framework assumes the prior noise as a standard Gaussian distribution, whereas the corresponding data distribution may be more complicated than th…

2022

Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token Embeddings

ACL 2022long

Recent studies have determined that the learned token embeddings of large-scale neural language models are degenerated to be anisotropic with a narrow-cone shape. This phenomenon, called the representation degeneration problem, facilitates an increase in the overall similarity between token embeddin…

Cited by 35SourcePDFScholar
2022

Stein Latent Optimization for Generative Adversarial Networks

ICLR 2022poster

Generative adversarial networks (GANs) with clustered latent spaces can perform conditional generation in a completely unsupervised manner. In the real world, the salient attributes of unlabeled data can be imbalanced. However, most of existing unsupervised conditional GANs cannot cluster attributes…

2022

Towards a Rigorous Evaluation of Time-Series Anomaly Detection

AAAI 2022technical

In recent years, proposed studies on time-series anomaly detection (TAD) report high F1 scores on benchmark TAD datasets, giving the impression of clear improvements in TAD. However, most studies apply a peculiar evaluation protocol called point adjustment (PA) before scoring. In this paper, we theo…

2022

Weakly Supervised Semantic Segmentation Using Out-of-Distribution Data

CVPR 2022poster

Weakly supervised semantic segmentation (WSSS) methods are often built on pixel-level localization maps obtained from a classifier. However, training on class labels only, classifiers suffer from the spurious correlation between foreground and background cues (e.g. train and rail), fundamentally bou…

Cited by 127PDFcodeScholar
2021

Accelerating Neural Architecture Search via Proxy Data

IJCAI 2021poster

Despite the increasing interest in neural architecture search (NAS), the significant computational cost of NAS is a hindrance to researchers. Hence, we propose to reduce the cost of NAS using proxy data, i.e., a representative subset of the target data, without sacrificing search performance. Even t…

2021

AligNART: Non-autoregressive Neural Machine Translation by Jointly Learning to Estimate Alignment and Translate

EMNLP 2021main

Non-autoregressive neural machine translation (NART) models suffer from the multi-modality problem which causes translation inconsistency such as token repetition. Most recent approaches have attempted to solve this problem by implicitly modeling dependencies between outputs. In this paper, we intro…

2021

Anti-Adversarially Manipulated Attributions for Weakly and Semi-Supervised Semantic Segmentation

CVPR 2021poster

Weakly supervised semantic segmentation produces a pixel-level localization from class labels; but a classifier trained on such labels is likely to restrict its focus to a small discriminative region of the target object. AdvCAM is an attribution map of an image that is manipulated to increase the c…

Cited by 307PDFcodeScholar
2021

BBAM: Bounding Box Attribution Map for Weakly Supervised Semantic and Instance Segmentation

CVPR 2021poster

Weakly supervised segmentation methods using bounding box annotations focus on obtaining a pixel-level mask from each box containing an object. Existing methods typically depend on a class-agnostic mask generator, which operates on the low-level information intrinsic to an image. In this work, we ut…

Cited by 232PDFcodeScholar
2021

ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models

ICCV 2021poster

Denoising diffusion probabilistic models (DDPM) have shown remarkable performance in unconditional image generation. However, due to the stochasticity of the generative process in DDPM, it is challenging to generate images with the desired semantics. In this work, we propose Iterative Latent Variabl…

Cited by 805PDFcodeScholar
2021

Reducing Information Bottleneck for Weakly Supervised Semantic Segmentation

NeurIPS 2021poster

Weakly supervised semantic segmentation produces pixel-level localization from class labels; however, a classifier trained on such labels is likely to focus on a small discriminative region of the target object. We interpret this phenomenon using the information bottleneck principle: the final layer…

2021

Removing Undesirable Feature Contributions Using Out-of-Distribution Data

ICLR 2021poster

Several data augmentation methods deploy unlabeled-in-distribution (UID) data to bridge the gap between the training and inference of neural networks. However, these methods have clear limitations in terms of availability of UID data and dependence of algorithms on pseudo-labels. Herein, we propose…

2021

XProtoNet: Diagnosis in Chest Radiography With Global and Local Explanations

CVPR 2021poster

Automated diagnosis using deep neural networks in chest radiography can help radiologists detect life-threatening diseases. However, existing methods only provide predictions without accurate explanations, undermining the trustworthiness of the diagnostic methods. Here, we present XProtoNet, a globa…

Cited by 144PDFScholar
2020

Adversarial Vertex Mixup: Toward Better Adversarially Robust Generalization

CVPR 2020oral

Adversarial examples cause neural networks to produce incorrect outputs with high confidence. Although adversarial training is one of the most effective forms of defense against adversarial examples, unfortunately, a large gap exists between test accuracy and training accuracy in adversarial trainin…

Cited by 151PDFcodeScholar
2020

Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search

NeurIPS 2020oral

Recently, text-to-speech (TTS) models such as FastSpeech and ParaNet have been proposed to generate mel-spectrograms from text in parallel. Despite the advantage, the parallel TTS models cannot be trained without guidance from autoregressive TTS models as their external aligners. In this work, we pr…

2020

NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity

NeurIPS 2020poster

Normalizing flows (NFs) have become a prominent method for deep generative models that allow for an analytic probability density estimation and efficient synthesis. However, a flow-based network is considered to be inefficient in parameter complexity because of reduced expressiveness of bijective ma…

2020

iCaps: An Interpretable Classifier via Disentangled Capsule Networks

ECCV 2020poster

We propose an interpretable Capsule Network, iCaps, for image classification. A capsule is a group of neurons nested inside each layer, and the one in the last layer is called a class capsule, which is a vector whose norm indicates a predicted probability for the class. Using the class capsule, exis…

Cited by 15SourcePDFScholar
2019

FickleNet: Weakly and Semi-Supervised Semantic Image Segmentation Using Stochastic Inference

CVPR 2019poster

The main obstacle to weakly supervised semantic image segmentation is the difficulty of obtaining pixel-level information from coarse image-level annotations. Most methods based on image-level annotations use localization maps obtained from the classifier, but these only focus on the small discrimin…

Cited by 557PDFScholar
2019

FloWaveNet : A Generative Flow for Raw Audio

ICML 2019oral

Most modern text-to-speech architectures use a WaveNet vocoder for synthesizing high-fidelity waveform audio, but there have been limitations, such as high inference time, in practical applications due to its ancestral sampling scheme. The recently suggested Parallel WaveNet and ClariNet has achieve…

2019

Frame-to-Frame Aggregation of Active Regions in Web Videos for Weakly Supervised Semantic Segmentation

ICCV 2019poster

When a deep neural network is trained on data with only image-level labeling, the regions activated in each image tend to identify only a small region of the target object. We propose a method of using videos automatically harvested from the web to identify a larger region of the target object by us…

Cited by 45PDFScholar
2017

Deep Recurrent Neural Network-Based Identification of Precursor microRNAs

NeurIPS 2017poster

MicroRNAs (miRNAs) are small non-coding ribonucleic acids (RNAs) which play key roles in post-transcriptional gene regulation. Direct identification of mature miRNAs is infeasible due to their short lengths, and researchers instead aim at identifying precursor miRNAs (pre-miRNAs). Many of the known…

2015

Boosted Categorical Restricted Boltzmann Machine for Computational Prediction of Splice Junctions

ICML 2015poster

Splicing refers to the elimination of non-coding regions in transcribed pre-messenger ribonucleic acid (RNA). Discovering splice sites is an important machine learning task that helps us not only to identify the basic units of genetic heredity but also to understand how different proteins are produc…

Cited by 93SourcePDFScholar