← Search

Joshua Tenenbaum

24 accepted papers

2026

Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind

ICRA 2026poster

Spoken language instructions are ubiquitous in agent collaboration. However, in real-world human-robot collaboration, following human spoken instructions can be challenging due to various speaker and environmental factors, such as background noise or mispronunciation. When faced with noisy auditory …

2024

MMToM-QA: Multimodal Theory of Mind Question Answering

ACL 2024long

Theory of Mind (ToM), the ability to understand people’s mental states, is an essential ingredient for developing machines with human-level social intelligence. Recent machine learning models, particularly large language models, seem to show some aspects of ToM understanding. However, existing ToM b…

2023

ConceptFusion: Open-set multimodal 3D mapping

RSS 2023poster

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-de…

2023

NeuSE: Neural SE(3)-Equivariant Embedding for Consistent Spatial Understanding with Objects

RSS 2023poster

We present NeuSE, a novel Neural SE(3)-Equivariant Embedding for objects, and illustrate how it supports object SLAM for consistent spatial understanding with longterm scene changes. NeuSE is a set of latent object embeddings created from partial object observations. It serves as a compact point clo…

Cited by 9SourcePDFScholar
2023

Tactile-Filter: Interactive Tactile Perception for Part Mating

RSS 2023poster

Humans rely on touch and tactile sensing for a lot of dexterous manipulation tasks. Our tactile sensing provides us with a lot of information regarding contact formations as well as geometric information about objects during any interaction. With this motivation, vision-based tactile sensors are bei…

2022

Discovering Generalizable Spatial Goal Representations via Graph-based Active Reward Learning

ICML 2022spotlight

In this work, we consider one-shot imitation learning for object rearrangement tasks, where an AI agent needs to watch a single expert demonstration and learn to perform the same task in different environments. To achieve a strong generalization, the AI agent must infer the spatial goal specificatio…

2022

Learning Iterative Reasoning through Energy Minimization

ICML 2022spotlight

Deep learning has excelled on complex pattern recognition tasks such as image classification and object recognition. However, it struggles with tasks requiring nontrivial reasoning, such as algorithmic computation. Humans are able to solve such tasks through iterative reasoning – spending more time…

2022

Planning with Diffusion for Flexible Behavior Synthesis

ICML 2022oral

Model-based reinforcement learning methods often use learning only for the purpose of recovering an approximate dynamics model, offloading the rest of the decision-making work to classical trajectory optimizers. While conceptually simple, this combination has a number of empirical shortcomings, sugg…

2022

Prompting Decision Transformer for Few-Shot Policy Generalization

ICML 2022spotlight

Human can leverage prior experience and learn novel tasks from a handful of demonstrations. In contrast to offline meta-reinforcement learning, which aims to achieve quick adaptation through better algorithm design, we investigate the effect of architecture inductive bias on the few-shot learning ca…

2021

A large-scale benchmark for few-shot program induction and synthesis

ICML 2021spotlight

A landmark challenge for AI is to learn flexible, powerful representations from small numbers of examples. On an important class of tasks, hypotheses in the form of programs provide extreme generalization capabilities from surprisingly few examples. However, whereas large natural few-shot learning i…

Cited by 24SourcePDFScholar
2021

AGENT: A Benchmark for Core Psychological Reasoning

ICML 2021spotlight

For machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants…

Cited by 96SourcePDFScholar
2021

Improved Contrastive Divergence Training of Energy-Based Models

ICML 2021spotlight

Contrastive divergence is a popular method of training energy-based models, but is known to have difficulties with training stability. We propose an adaptation to improve contrastive divergence training by scrutinizing a gradient term that is difficult to calculate and is often left out for convenie…

Cited by 169SourcePDFScholar
2021

Learning Symbolic Operators for Task and Motion Planning

IROS 2021poster

Robotic planning problems in hybrid state and action spaces can be solved by integrated task and motion planners (TAMP) that handle the complex interaction between motion-level decisions and task-level plan feasibility. TAMP approaches rely on domain-specific symbolic operators to guide the task-lev…

Cited by 108SourceScholar
2021

Leveraging Language to Learn Program Abstractions and Search Heuristics

ICML 2021spotlight

Inductive program synthesis, or inferring programs from examples of desired behavior, offers a general paradigm for building interpretable, robust, andgeneralizable machine learning systems. Effective program synthesis depends on two key ingredients: a strong library of functions from which to build…

Cited by 69SourcePDFScholar
2020

A Long Horizon Planning Framework for Manipulating Rigid Pointcloud Objects

CoRL 2020

We present a framework for solving long-horizon planning problems involving manipulation of rigid objects that operates directly from a point-cloud observation. Our method plans in the space of object subgoals and frees the planner from reasoning about robot-object interaction dynamics. We show that

2020

Visual Grounding of Learned Physical Models

ICML 2020poster

Humans intuitively recognize objects’ physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environments, while intrinsic to humans, remain challenging to state-of-the-art computati…

Cited by 82SourcePDFScholar
2019

DensePhysNet: Learning Dense Physical Object Representations Via Multi-Step Dynamic Interactions

RSS 2019poster

We study the problem of learning physical object representations for robot manipulation. Understanding object physics is critical for successful object manipulation, but also challenging because physical object properties can rarely be inferred from the object's static appearance. In this paper, we…

Cited by 125SourcePDFScholar
2019

Entity Abstraction in Visual Model-Based Reinforcement Learning

CoRL 2019

We present OP3, a framework for model-based reinforcement learning that acquires object representations from raw visual observations without supervision and uses them to predict and plan. To ground these abstract representations of entities to actual objects in the world, we formulate an interactive

2018

Differentiable Physics and Stable Modes for Tool-Use and Manipulation Planning

RSS 2018poster

We consider the problem of sequential manipulation and tool-use planning in domains that include physical interactions such as hitting and throwing. The approach integrates a Task And Motion Planning formulation with primitives that either impose stable kinematic constraints or differentiabl…

2017

A Compositional Object-Based Approach to Learning Physical Dynamics

ICLR 2017poster

We present the Neural Physics Engine (NPE), a framework for learning simulators of intuitive physics that naturally generalize across variable object count and different scene configurations. We propose a factorization of a physical scene into composable object-based representations and a neural net…

Cited by 530SourceScholar