← Search

Smitha Milli

8 accepted papers

2026

Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset

ICLR 2026poster

How can large language models (LLMs) serve users with varying preferences that may conflict across cultural, political, or other dimensions? To advance this challenge, this paper establishes four key results. First, we demonstrate, through a large-scale multilingual human study with representative s…

Cited by 0SourcecodeScholar
2026

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

ICLR 2026oral

Preference data is widely used for aligning language models, but remains largely opaque. While prior work has studied specific aspects of annotator preference (e.g., length or sycophancy), automatically inferring preferences without pre-specifying hypotheses remains challenging. We introduce *What's…

Cited by 0SourcecodeScholar
2025

Position: Democratic AI is Possible. The Democracy Levels Framework Shows How It Might Work.

ICML 2025poster

This position paper argues that effectively "democratizing AI" requires democratic governance and alignment of AI, and that this is particularly valuable for decisions with systemic societal impacts. Initial steps—such as Meta's *Community Forums* and Anthropic's *Collective Constitutional AI*—have…

Cited by 0SourcePDFScholar
2025

Representative Ranking for Deliberation in the Public Sphere

ICML 2025poster

Online comment sections, such as those on news sites or social media, have the potential to foster informal public deliberation, However, this potential is often undermined by the frequency of toxic or low-quality exchanges that occur in these settings. To combat this, platforms increasingly leverag…

Cited by 0SourcePDFScholar
2020

Reward-rational (implicit) choice: A unifying formalism for reward learning

NeurIPS 2020poster

It is often difficult to hand-specify what the correct reward function is for a task, so researchers have instead aimed to learn reward functions from human behavior or feedback. The types of behavior interpreted as evidence of the reward function have expanded greatly in recent years. We've gone fr…

Cited by 226SourcePDFScholar
2019

Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning

UAI 2019poster

It is incredibly easy for a system designer to misspecify the objective for an autonomous system (“robot"), thus motivating the desire to have the robot learn the objective from human behavior instead. Recent work has suggested that people have an interest in the robot performing well, and will thu…

Cited by 26SourcePDFScholar