← Search

Tobias Gerstenberg

7 accepted papers

2026

Spot The Ball: A Benchmark for Visual Social Inference

CVPR 2026

Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This capacity drives everyday social reasoning in humans and is critical for developing more human-like AI agents. We introduce Spo

Cited by 0SourcecodeScholar
2025

Causal-PIK: Causality-based Physical Reasoning with a Physics-Informed Kernel

ICML 2025poster

Tasks that involve complex interactions between objects with unknown dynamics make planning before execution difficult. These tasks require agents to iteratively improve their actions after actively exploring causes and effects in the environment. For these type of tasks, we propose Causal-PIK, a me…

Cited by 0SourcePDFScholar
2024

MARPLE: A Benchmark for Long-Horizon Inference

NeurIPS 2024poster

Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchm…

2024

Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels

NeurIPS 2024poster

When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, t…

2023

MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment Tasks

NeurIPS 2023poster

Human commonsense understanding of the physical and social world is organized around intuitive theories. These theories support making causal and moral judgments. When something bad happens, we naturally ask: who did what, and why? A rich literature in cognitive science has studied people's causal a…

Cited by 40SourcePDFScholar
2023

Understanding Social Reasoning in Language Models with Language Models

NeurIPS 2023spotlight

As Large Language Models (LLMs) become increasingly integrated into our everyday lives, understanding their ability to comprehend human mental states becomes critical for ensuring effective interactions. However, despite the recent attempts to assess the Theory-of-Mind (ToM) reasoning capabilities o…

Cited by 123SourcePDFScholar
2022

Uncalibrated Models Can Improve Human-AI Collaboration

NeurIPS 2022accept

In many practical applications of AI, an AI model is used as a decision aid for human users. The AI provides advice that a human (sometimes) incorporates into their decision-making process. The AI advice is often presented with some measure of "confidence" that the human can use to calibrate how muc…