← Search

Changho Shin

7 accepted papers

2026

CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation

ICML 2026poster

LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared l…

Cited by 0SourceScholar
2024

OTTER: Effortless Label Distribution Adaptation of Zero-shot Models

NeurIPS 2024poster

Popular zero-shot models suffer due to artifacts inherited from pretraining. One particularly detrimental issue, caused by unbalanced web-scale pretraining data, is mismatched label distribution. Existing approaches that seek to repair the label distribution are not suitable in zero-shot settings, a…

2023

Mitigating Source Bias for Fairer Weak Supervision

NeurIPS 2023poster

Weak supervision enables efficient development of training sets by reducing the need for ground truth labels. However, the techniques that make weak supervision attractive---such as integrating any source of signal to estimate unknown labels---also entail the danger that the produced pseudolabels ar…

2022

Universalizing Weak Supervision

ICLR 2022poster

Weak supervision (WS) frameworks are a popular way to bypass hand-labeling large datasets for training data-hungry models. These approaches synthesize multiple noisy but cheaply-acquired estimates of labels into a set of high-quality pseudo-labels for downstream training. However, the synthesis tech…

Cited by 43SourcePDFScholar