← Search

Ramya Korlakai Vinayak

15 accepted papers

2025

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

ICCV 2025poster

Popular text-to-image (T2I) systems are trained on web-scraped data, which is heavily Amero and Euro-centric, underrepresenting the cultures of the Global South. To analyze these biases, we introduce CuRe, a novel and scalable benchmarking and scoring suite for cultural representativeness that lever…

2025

PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment

ICLR 2025poster

Foundation models trained on internet-scale data benefit from extensive alignment to human preferences before deployment. However, existing methods typically assume a homogeneous preference shared by all individuals, overlooking the diversity inherent in human values. In this work, we propose a gene…

Cited by 2SourcePDFScholar
2025

Rethinking Confidence Scores and Thresholds in Pseudolabeling-based SSL

ICML 2025poster

Modern semi-supervised learning (SSL) methods rely on pseudolabeling and consistency regularization. Pseudolabeling is typically performed by comparing the model's confidence scores and a predefined threshold. While several heuristics have been proposed to improve threshold selection, the underlyin…

Cited by 0SourcePDFScholar
2024

Learning Populations of Preferences via Pairwise Comparison Queries

AISTATS 2024poster

Ideal point based preference learning using pairwise comparisons of type "Do you prefer a or b?" has emerged as a powerful tool for understanding how we make preferences. Existing preference learning approaches assume homogeneity and focus on learning preference on average over the population or req…

Cited by 5SourcePDFScholar
2024

Limitations of Face Image Generation

AAAI 2024technical

Text-to-image diffusion models have achieved widespread popularity due to their unprecedented image generation capability. In particular, their ability to synthesize and modify human faces has spurred research into using generated face images in both training data augmentation and model performance…

2024

Pearls from Pebbles: Improved Confidence Functions for Auto-labeling

NeurIPS 2024poster

Auto-labeling is an important family of techniques that produce labeled training sets with minimum manual annotation. A prominent variant, threshold-based auto-labeling (TBAL), works by finding thresholds on a model's confidence scores above which it can accurately automatically label unlabeled data…

Cited by 2SourcePDFScholar
2024

Taming False Positives in Out-of-Distribution Detection with Human Feedback

AISTATS 2024poster

Robustness to out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models in the open world. Recent works have focused on designing scoring functions to quantify OOD uncertainty. Setting appropriate thresholds for these scoring functions for OOD detection is chall…

2023

Promises and Pitfalls of Threshold-based Auto-labeling

NeurIPS 2023spotlight

Creating large-scale high-quality labeled datasets is a major bottleneck in supervised machine learning workflows. Threshold-based auto-labeling (TBAL), where validation data obtained from humans is used to find a confidence threshold above which the data is machine-labeled, reduces reliance on manu…

2022

One for All: Simultaneous Metric and Preference Learning over Multiple Users

NeurIPS 2022accept

This paper investigates simultaneous preference and metric learning from a crowd of respondents. A set of items represented by $d$-dimensional feature vectors and paired comparisons of the form ``item $i$ is preferable to item $j$'' made by each user is given. Our model jointly learns a distance met…

2020

Estimating the Number and Effect Sizes of Non-null Hypotheses

ICML 2020poster

We study the problem of estimating the distribution of effect sizes (the mean of the test statistic under the alternate hypothesis) in a multiple testing setting. Knowing this distribution allows us to calculate the power (type II error) of any experimental design. We show that it is possible to est…

2019

Maximum Likelihood Estimation for Learning Populations of Parameters

ICML 2019oral

Consider a setting with $N$ independent individuals, each with an unknown parameter, $p_i \in [0, 1]$ drawn from some unknown distribution $P^\star$. After observing the outcomes of $t$ independent Bernoulli trials, i.e., $X_i \sim \text{Binomial}(t, p_i)$ per individual, our objective is to accurat…

Cited by 51SourcePDFScholar