← Search

Murray Campbell

11 accepted papers

2025

Position: Theory of Mind Benchmarks are Broken for Large Language Models

ICML 2025poster

Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks for LLMs are overwhelmingly inspired by the methods used to t…

Cited by 0SourcePDFScholar
2024

On the generalization capacity of neural networks during generic multimodal reasoning

ICLR 2024poster

The advent of the Transformer has led to the development of large language models (LLM), which appear to demonstrate human-like capabilities. To assess the generality of this class of models and a variety of other base neural network architectures to multimodal domains, we evaluated and compared the…

2023

JECC: Commonsense Reasoning Tasks Derived from Interactive Fictions

ACL 2023findings

Commonsense reasoning simulates the human ability to make presumptions about our physical world, and it is an essential cornerstone in building general AI systems. We proposea new commonsense reasoning dataset based on human’s Interactive Fiction (IF) gameplaywalkthroughs as human players demonstrat…

2021

Circles are like Ellipses, or Ellipses are like Circles? Measuring the Degree of Asymmetry of Static and Contextual Word Embeddings and the Implications to Representation Learning

AAAI 2021technical

Human judgments of word similarity have been a popular method of evaluating the quality of word embedding. But it fails to measure the geometry properties such as asymmetry. For example, it is more natural to say ``Ellipses are like Circles'' than ``Circles are like Ellipses''. Such asymmetry has be…

Cited by 0SourcePDFScholar
2021

Efficient Text-based Reinforcement Learning by Jointly Leveraging State and Commonsense Graph Representations

ACL 2021short

Text-based games (TBGs) have emerged as useful benchmarks for evaluating progress at the intersection of grounded language understanding and reinforcement learning (RL). Recent work has proposed the use of external knowledge to improve the efficiency of RL agents for TBGs. In this paper, we posit th…

Cited by 17SourcePDFScholar
2021

Mental Models of AI Agents in a Cooperative Game Setting (Extended Abstract)

IJCAI 2021poster

As more and more forms of AI become prevalent, it becomes increasingly important to understand how people develop mental models of these systems. In this work we study people's mental models of an AI agent in a cooperative word guessing game. We run a study in which people play the game with an AI a…

Cited by 0SourcePDFScholar
2021

Text-based RL Agents with Commonsense Knowledge: New Challenges, Environments and Baselines

AAAI 2021technical

Text-based games have emerged as an important test-bed for Reinforcement Learning (RL) research, requiring RL agents to combine grounded language understanding with sequential decision making. In this paper, we examine the problem of infusing RL agents with commonsense knowledge. Such knowledge woul…

2018

Eigenoption Discovery through the Deep Successor Representation

ICLR 2018poster

Options in reinforcement learning allow agents to hierarchically decompose a task into subtasks, having the potential to speed up learning and planning. However, autonomously learning effective sets of options is still a major challenge in the field. In this paper we focus on the recently introduced…

Cited by 194SourcePDFScholar
2018

Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering

ICLR 2018poster

Very recently, it comes to be a popular approach for answering open-domain questions by first searching question-related passages, then applying reading comprehension models to extract answers. Existing works usually extract answers from single passages independently, thus not fully make use of the…

2017

Learning to Query, Reason, and Answer Questions On Ambiguous Texts

ICLR 2017poster

A key goal of research in conversational systems is to train an interactive agent to help a user with a task. Human conversation, however, is notoriously incomplete, ambiguous, and full of extraneous detail. To operate effectively, the agent must not only understand what was explicitly conveyed but…

Cited by 31SourceScholar