← Search

Jan-Philipp Fränken

4 accepted papers

2025

PERSONA: A Reproducible Testbed for Pluralistic Alignment

COLING 2025main

The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the plurality of user opinions, instead reinforcing majority viewpoints and marginalizing minority perspectives. We introdu…

Cited by 12SourcePDFScholar
2024

MARPLE: A Benchmark for Long-Horizon Inference

NeurIPS 2024poster

Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchm…

2024

Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels

NeurIPS 2024poster

When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, t…

2023

Understanding Social Reasoning in Language Models with Language Models

NeurIPS 2023spotlight

As Large Language Models (LLMs) become increasingly integrated into our everyday lives, understanding their ability to comprehend human mental states becomes critical for ensuring effective interactions. However, despite the recent attempts to assess the Theory-of-Mind (ToM) reasoning capabilities o…

Cited by 123SourcePDFScholar