← Search

Hye Won Chung

16 accepted papers

2025

CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder

NeurIPS 2025poster

Multimodal dataset distillation aims to synthesize a small set of image-text pairs that enables efficient training of large-scale vision-language models. While dataset distillation has shown promise in unimodal tasks, extending it to multimodal contrastive learning presents key challenges: learning…

Cited by 0SourceScholar
2025

Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels

ICLR 2025poster

We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply. Our theoretical analysis reveals that multi-round self-distillation effectively…

Cited by 0SourcePDFScholar
2025

SNAP: Low-Latency Test-Time Adaptation with Sparse Updates

NeurIPS 2025poster

Test-Time Adaptation (TTA) adjusts models using unlabeled test data to handle dynamic distribution shifts. However, existing methods rely on frequent adaptation and high computational cost, making them unsuitable for resource-constrained edge environments. To address this, we propose SNAP, a sparse…

Cited by 0SourcecodeScholar
2024

BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad Ranges

ICML 2024poster

Data subset selection aims to find a smaller yet informative subset of a large dataset that can approximate the full-dataset training, addressing challenges associated with training neural networks on large-scale datasets. However, existing methods tend to specialize in either high or low selection…

2024

SelMatch: Effectively Scaling Up Dataset Distillation via Selection-Based Initialization and Partial Updates by Trajectory Matching

ICML 2024poster

Dataset distillation aims to synthesize a small number of images per class (IPC) from a large dataset to approximate full dataset training with minimal performance loss. While effective in very small IPC ranges, many distillation methods become less effective, even underperforming random sample sele…

Cited by 9SourcePDFScholar
2023

Efficient Algorithms for Exact Graph Matching on Correlated Stochastic Block Models with Constant Correlation

ICML 2023poster

We consider the problem of graph matching, or learning vertex correspondence, between two correlated stochastic block models (SBMs). The graph matching problem arises in various fields, including computer vision, natural language processing and bioinformatics, and in particular, matching graphs with…

2023

Recovering Top-Two Answers and Confusion Probability in Multi-Choice Crowdsourcing

ICML 2023poster

Crowdsourcing has emerged as an effective platform for labeling large amounts of data in a cost- and time-efficient manner. Most previous work has focused on designing an efficient algorithm to recover only the ground-truth labels of the data. In this paper, we consider multi-choice crowdsourcing ta…

2023

Test-Time Adaptation via Self-Training with Nearest Neighbor Information

ICLR 2023poster

Test-time adaptation (TTA) aims to adapt a trained classifier using online unlabeled test data only, without any information related to the training procedure. Most existing TTA methods adapt the trained classifier using the classifier's prediction on the test data as pseudo-label. However, under te…

2021

Self-Diagnosing GAN: Diagnosing Underrepresented Samples in Generative Adversarial Networks

NeurIPS 2021poster

Despite remarkable performance in producing realistic samples, Generative Adversarial Networks (GANs) often produce low-quality samples near low-density regions of the data manifold, e.g., samples of minor groups. Many techniques have been developed to improve the quality of generated samples, eithe…

2018

Unequal Error Protection Querying Policies for the Noisy 20 Questions Problem

ICASSP 2018accepted

We propose a non-adaptive unequal error protection (UEP) querying policy based on superposition coding for the noisy 20 questions problem. In this problem, a player wishes to successively refine an estimate of the value of a continuous random variable by posing binary queries and receiving noisy res…

Cited by 0SourceScholar