← Search

Chris Paxton

52 accepted papers

2025

Dynamem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation

ICRA 2025

Significant progress has been made in openvocabulary mobile manipulation, where the goal is for a robot to perform tasks in any environment given a natural language description. However, most current systems assume a static environment, which limits the system's applicability in realworld scenarios

Cited by 34SourcecodeScholar
2025

GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering

CoRL 2025poster

In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment in order to answer a situated question with confidence. This remains a challenging problem in robotics, due to the difficulties in obtaining useful semantic representations, updati…

Cited by 0SourcecodeScholar
2025

Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments

ICRA 2025

Robot models, particularly those trained with large amounts of data, have recently shown a plethora of real-world manipulation and navigation capabilities. Several independent efforts have shown that given sufficient training data in an environment, robot policies can generalize to demonstrated vari

Cited by 49SourcecodeScholar
2024

Demonstrating OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

RSS 2024poster

Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effectively control mobile systems, and grasping models that can handle a wide range of o…

2024

GOAT: GO to Any Thing

RSS 2024poster

In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system cap…

2024

HACMan++: Spatially-Grounded Motion Primitives for Manipulation

RSS 2024poster

Although end-to-end robot learning has shown some success for robot manipulation, the learned policies are often not sufficiently robust to variations in object pose or geometry. To improve the policy generalization, we introduce spatially-grounded parameterized motion primitives in our method HACMa…

2024

OpenEQA: Embodied Question Answering in the Era of Foundation Models

CVPR 2024poster

We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory exemplified by agents on smart glasses or b…

Cited by 118SourcePDFScholar
2023

CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

RSS 2023poster

We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this…

2023

HACMan: Learning Hybrid Actor-Critic Maps for 6D Non-Prehensile Manipulation

CoRL 2023oral

Manipulating objects without grasping them is an essential component of human dexterity, referred to as non-prehensile manipulation. Non-prehensile manipulation may enable more complex interactions with the objects, but also presents challenges in reasoning about gripper-object interactions. In this…

Cited by 22SourcecodeScholar
2023

HomeRobot: Open-Vocabulary Mobile Manipulation

CoRL 2023poster

HomeRobot (noun): An affordable compliant robot that navigates homes and manipulates a wide range of objects in order to complete everyday tasks. Open-Vocabulary Mobile Manipulation (OVMM) is the problem of picking any object in any unseen environment, and placing it in a commanded location. This i…

Cited by 98SourcecodeScholar
2023

Navigating to Objects Specified by Images

ICCV 2023poster

Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform this task in both simulation and the real world. Our modula…

Cited by 41PDFScholar
2023

SLAP: Spatial-Language Attention Policies

CoRL 2023poster

Despite great strides in language-guided manipulation, existing work has been constrained to table-top settings. Table-tops allow for perfect and consistent camera angles, properties are that do not hold in mobile manipulation. Task plans that involve moving around the environment must be robust to…

Cited by 8SourcecodeScholar
2023

StructDiffusion: Language-Guided Creation of Physically-Valid Structures using Unseen Objects

RSS 2023poster

Robots operating in human environments must be able to rearrange objects into semantically-meaningful configurations, even if these objects are previously unseen. In this work, we focus on the problem of building physically-valid structures without step-by-step instructions. We propose StructDiffusi…

Cited by 45SourcePDFScholar
2023

Task and Motion Planning with Large Language Models for Object Rearrangement

IROS 2023poster

Multi-object rearrangement is a crucial skill for service robots, and commonsense reasoning is frequently needed in this process. However, achieving commonsense arrangements requires knowledge about objects, which is hard to transfer to robots. Large language models (LLMs) are one potential source o…

Cited by 192SourceScholar
2023

USA-Net: Unified Semantic and Affordance Representations for Robot Memory

IROS 2023poster

In order for robots to follow open-ended instructions like “go open the brown cabinet over the sink,” they require an understanding of both the scene geometry and the semantics of their environment. Robotic systems often handle these through separate pipelines, sometimes using very different represe…

Cited by 13SourceScholar
2022

Correcting Robot Plans with Natural Language Feedback

RSS 2022poster

When humans design cost or goal specifications for robots, they often produce specifications that are ambiguous, under-specified, or beyond planners’ ability to solve. In these cases, corrections provide a valuable tool for human-in-the-loop robot control. Corrections might take the form of new goal…

Cited by 110SourcePDFScholar
2022

HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers

ICRA 2022poster

We introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics…

Cited by 29SourcecodeScholar
2022

IFOR: Iterative Flow Minimization for Robotic Object Rearrangement

CVPR 2022poster

Accurate object rearrangement from vision is a crucial problem for a wide variety of real-world robotics applications in unstructured environments. We propose IFOR, Iterative Flow Minimization for Robotic Object Rearrangement, an end-to-end method for the challenging problem of object rearrangement…

Cited by 59PDFcodeScholar
2022

Learning Perceptual Concepts by Bootstrapping From Human Queries

RA-L 2022

When robots operate in human environments, it's critical that humans can quickly teach them new <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">concepts:</i> object-centric properties of the environment that they care about (e.g., objects <italic xml

Cited by 17SourceScholar
2022

Model Predictive Control for Fluid Human-to-Robot Handovers

ICRA 2022poster

Human-robot handover is a fundamental yet challenging task in human-robot interaction and collaboration. Recently, remarkable progressions have been made in human-to-robot handovers of unknown objects by using learning-based grasp generators. However, how to responsively generate smooth motions to t…

Cited by 31SourceScholar
2022

Pre-Trained Language Models for Interactive Decision-Making

NeurIPS 2022accept

Language model (LM) pre-training is useful in many language processing tasks. But can pre-trained LMs be further leveraged for more general machine learning problems? We propose an approach for using LMs to scaffold learning and generalization in general sequential decision-making problems. In this…

Cited by 229SourcePDFScholar
2022

StructFormer: Learning Spatial Structure for Language-Guided Semantic Rearrangement of Novel Objects

ICRA 2022poster

Geometric organization of objects into semantically meaningful arrangements pervades the built world. As such, assistive robots operating in warehouses, offices, and homes would greatly benefit from the ability to recognize and rearrange objects into these semantically meaningful structures. To be u…

Cited by 104SourceScholar
2022

Transporters with Visual Foresight for Solving Unseen Rearrangement Tasks

IROS 2022poster

Rearrangement tasks have been identified as a crucial challenge for intelligent robotic manipulation, but few methods allow for precise construction of unseen structures. We propose a visual foresight model for pick-and-place rearrangement manipulation which is able to learn efficiently. In addition…

Cited by 16SourcecodeScholar
2021

A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution

CoRL 2021poster

Natural language provides an accessible and expressive interface to specify long-term tasks for robotic agents. However, non-experts are likely to specify such tasks with high-level instructions, which abstract over specific robot actions through several layers of abstraction. We propose that key to…

Cited by 151SourcecodeScholar
2021

Alternative Paths Planner (APP) for Provably Fixed-time Manipulation Planning in Semi-structured Environments

ICRA 2021poster

In many applications, including logistics and manufacturing, robot manipulators operate in semi-structured environments alongside humans or other robots. These environments are largely static, but they may contain some movable obstacles that the robot must avoid. Manipulation tasks in these applicat…

Cited by 8SourceScholar
2021

Automated Generation of Robotic Planning Domains from Observations

IROS 2021poster

Automated planning enables robots to find plans to achieve complex, long-horizon tasks, given a planning domain. This planning domain consists of a list of actions, with their associated preconditions and effects, and is usually manually defined by a human expert, which is very time-consuming or eve…

Cited by 46SourceScholar
2021

NeRP: Neural Rearrangement Planning for Unknown Objects

RSS 2021poster

Robots will be expected to manipulate a wide variety of objects in complex and arbitrary ways as they become more widely used in human environments. As such; the rearrangement of objects has been noted to be an important benchmark for AI capabilities in recent years. We propose NeRP (Neural Rearrang…

Cited by 84SourcePDFScholar
2021

Predicting Stable Configurations for Semantic Placement of Novel Objects

CoRL 2021poster

Human environments contain numerous objects configured in a variety of arrangements. Our goal is to enable robots to repose previously unseen objects according to learned semantic relationships in novel environments. We break this problem down into two parts: (1) finding physically valid locations f…

Cited by 54SourcecodeScholar
2021

Reactive Human-to-Robot Handovers of Arbitrary Objects

ICRA 2021poster

Human-robot object handovers have been an actively studied area of robotics over the past decade; however, very few techniques and systems have addressed the challenge of handing over diverse objects with arbitrary appearance, size, shape, and deformability. In this paper, we present a vision-based…

Cited by 94SourceScholar
2021

Reactive Long Horizon Task Execution via Visual Skill and Precondition Models

IROS 2021poster

Zero-shot execution of unseen robotic tasks is important to allowing robots to perform a wide variety of tasks in human environments, but collecting the amounts of data necessary to train end-to-end policies in the real-world is often infeasible. We describe an approach for sim-to-real training that…

Cited by 21SourceScholar
2021

SORNet: Spatial Object-Centric Representations for Sequential Manipulation

CoRL 2021oral

Sequential manipulation tasks require a robot to perceive the state of an environment and plan a sequence of actions leading to a desired goal state, where the ability to reason about spatial relationships among object entities from raw sensor inputs is crucial. Prior works relying on explicit state…

Cited by 86SourcecodeScholar
2020

"Good Robot!": Efficient Reinforcement Learning for Multi-Step Visual Tasks with Sim to Real Transfer

RA-L 2020

Current Reinforcement Learning (RL) algorithms struggle with long-horizon tasks where time can be wasted exploring dead ends and task progress may be easily reversed. We develop the SPOT framework, which explores within action safety zones, learns about unsafe regions without exploring them, and pri

Cited by 73SourcecodeScholar
2020

6-DOF Grasping for Target-driven Object Manipulation in Clutter

ICRA 2020poster

Grasping in cluttered environments is a fundamental but challenging robotic skill. It requires both reasoning about unseen object parts and potential collisions with the manipulator. Most existing data-driven approaches avoid this problem by limiting themselves to top-down planar grasps which is ins…

Cited by 263SourcecodeScholar
2020

Collaborative Interaction Models for Optimized Human-Robot Teamwork

IROS 2020poster

Effective human-robot collaboration requires informed anticipation. The robot must anticipate the human’s actions, but also react quickly and intuitively when its predictions are wrong. The robot must plan its actions to account for the human’s own plan, with the knowledge that the human’s behavior…

Cited by 22SourceScholar
2020

Motion Reasoning for Goal-Based Imitation Learning

ICRA 2020poster

We address goal-based imitation learning, where the aim is to output the symbolic goal from a third-person video demonstration. This enables the robot to plan for execution and reproduce the same goal in a completely different environment. The key challenge is that the goal of a video demonstration…

Cited by 20SourceScholar
2020

Online Replanning in Belief Space for Partially Observable Task and Motion Problems

ICRA 2020poster

To solve multi-step manipulation tasks in the real world, an autonomous robot must take actions to observe its environment and react to unexpected observations. This may require opening a drawer to observe its contents or moving an object out of the way to examine the space behind it. Upon receiving…

Cited by 145SourcecodeScholar
2020

Transferable Task Execution from Pixels through Deep Planning Domain Learning

ICRA 2020poster

While robots can learn models to solve many manipulation tasks from raw visual input, they cannot usually use these models to solve new problems. On the other hand, symbolic planning methods such as STRIPS have long been able to solve new problems given only a domain definition and a symbolic goal,…

Cited by 52SourceScholar
2019

Prospection: Interpretable plans from language by predicting the future

ICRA 2019poster

High-level human instructions often correspond to behaviors with multiple implicit steps. In order for robots to be useful in the real world, they must be able to to reason over both motions and intermediate goals implied by human instructions. In this work, we propose a framework for learning repre…

Cited by 58SourceScholar
2019

Representing Robot Task Plans as Robust Logical-Dynamical Systems

IROS 2019poster

It is difficult to create robust, reusable, and reactive behaviors for robots that can be easily extended and combined. Frameworks such as Behavior Trees are flexible but difficult to characterize, especially when designing reactions and recovery behaviors to consistently converge to a desired goal…

Cited by 82SourceScholar
2019

The CoSTAR Block Stacking Dataset: Learning with Workspace Constraints

IROS 2019poster

A robot can now grasp an object more effectively than ever before, but once it has the object what happens next? We show that a mild relaxation of the task and workspace constraints implicit in existing object grasping datasets can cause neural network based grasping algorithms to fail on even a sim…

Cited by 12SourceScholar
2019

Uncertainty-Aware Occupancy Map Prediction Using Generative Networks for Robot Navigation

ICRA 2019poster

Efficient exploration through unknown environments remains a challenging problem for robotic systems. In these situations, the robot's ability to reason about its future motion is often severely limited by sensor field of view (FOV). By contrast, biological systems routinely make decisions by taking…

Cited by 67SourceScholar
2018

Evaluating Methods for End-User Creation of Robot Task Plans

IROS 2018poster

How can we enable users to create effective, perception-driven task plans for collaborative robots? We conducted a 35-person user study with the Behavior Tree-based CoSTAR system to determine which strategies for end user creation of generalizable robot task plans are most usable and effctive. CoSTA…

Cited by 47SourceScholar
2017

CoSTAR: Instructing collaborative robots with behavior trees and vision

ICRA 2017poster

For collaborative robots to become useful, end users who are not robotics experts must be able to instruct them to perform a variety of tasks. With this goal in mind, we developed a system for end-user creation of robust task plans with a broad range of capabilities. CoSTAR: the Collaborative System…

Cited by 226SourcecodeScholar
2017

Combining neural networks and tree search for task and motion planning in challenging environments

IROS 2017poster

Task and motion planning subject to Linear Temporal Logic (LTL) specifications in complex, dynamic environments requires efficient exploration of many possible future worlds. Model-free reinforcement learning has proven successful in a number of challenging tasks, but shows poor performance on tasks…

Cited by 153SourceScholar
2016

Do what i want, not what i did: Imitation of skills by planning sequences of actions

IROS 2016poster

We propose a learning-from-demonstration approach for grounding actions from expert data and an algorithm for using these actions to perform a task in new environments. Our approach is based on an application of sampling-based motion planning to search through the tree of discrete, high-level action…

Cited by 28SourceScholar
2015

A framework for end-user instruction of a robot assistant for manufacturing

ICRA 2015poster

Small Manufacturing Entities (SMEs) have not incorporated robotic automation as readily as large companies due to rapidly changing product lines, complex and dexterous tasks, and the high cost of start-up. While recent low-cost robots such as the Universal Robots UR5 and Rethink Robotics Baxter are…

Cited by 144SourceScholar
2015

An incremental approach to learning generalizable robot tasks from human demonstration

ICRA 2015poster

Dynamic Movement Primitives (DMPs) are a common method for learning a control policy for a task from demonstration. This control policy consists of differential equations that can create a smooth trajectory to a new goal point. However, DMPs only have a limited ability to generalize the demonstratio…

Cited by 56SourceScholar