← Search

Kanishk Gandhi

5 accepted papers

2024

Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels

NeurIPS 2024poster

When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, t…

2023

Understanding Social Reasoning in Language Models with Language Models

NeurIPS 2023spotlight

As Large Language Models (LLMs) become increasingly integrated into our everyday lives, understanding their ability to comprehend human mental states becomes critical for ensuring effective interactions. However, despite the recent attempts to assess the Theory-of-Mind (ToM) reasoning capabilities o…

Cited by 123SourcePDFScholar
2022

Eliciting Compatible Demonstrations for Multi-Human Imitation Learning

CoRL 2022poster

Imitation learning from human-provided demonstrations is a strong approach for learning policies for robot manipulation. While the ideal dataset for imitation learning is homogenous and low-variance - reflecting a single, optimal method for performing a task - natural human behavior has a great deal…

Cited by 22SourceScholar
2021

Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of others

NeurIPS 2021poster

To achieve human-like common sense about everyday life, machine learning systems must understand and reason about the goals, preferences, and actions of other agents in the environment. By the end of their first year of life, human infants intuitively achieve such common sense, and these cognitive a…

Cited by 69SourcePDFScholar