← Search

Michiel A. Bakker

5 accepted papers

2026

Benchmarking Overton Pluralism in LLMs

ICLR 2026poster

We introduce a novel framework for measuring Overton pluralism in LLMs—the extent to which diverse viewpoints are represented in model outputs. We (i) formalize Overton pluralism as a set-coverage metric (OVERTONSCORE), (ii) conduct a large-scale US-representative human study (N=1209; 60 questions;…

Cited by 0SourcecodeScholar
2026

Robust Preference Optimization: Aligning Language Models with Noisy Preference Feedback

ICLR 2026poster

Standard human preference-based alignment methods, such as Reinforcement Learning from Human Feedback (RLHF), are a cornerstone technology for aligning Large Language Models (LLMs) with human values. However, these methods are all underpinned by a strong assumption that the collected preference data…

Cited by 0SourceScholar
2025

Position: Democratic AI is Possible. The Democracy Levels Framework Shows How It Might Work.

ICML 2025poster

This position paper argues that effectively "democratizing AI" requires democratic governance and alignment of AI, and that this is particularly valuable for decisions with systemic societal impacts. Initial steps—such as Meta's *Community Forums* and Anthropic's *Collective Constitutional AI*—have…

Cited by 0SourcePDFScholar
2025

Value Profiles for Encoding Human Variation

EMNLP 2025

Modelling human variation in rating tasks is crucial for enabling AI systems for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using value profiles – natural language descriptions of underlying values compressed from in-context de

Cited by 0SourcePDFScholar
2022

Fine-tuning language models to find agreement among humans with diverse preferences

NeurIPS 2022accept

Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static and homogeneous across individuals, so that aligning to a single "generic" user will confer more general alignment. Her…

Cited by 258SourcePDFScholar