← Search

Soumith Chintala

15 accepted papers

2026

Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models

RSS 2026poster

The prevalent paradigm in robot learning attempts to generalize across environments, embodiments, and tasks with language prompts at runtime. A fundamental tension limits this approach: language is often too abstract to guide the concrete physical understanding required for robust manipulation. In t…

Cited by 0SourceScholar
2025

Dynamem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation

ICRA 2025

Significant progress has been made in openvocabulary mobile manipulation, where the goal is for a robot to perform tasks in any environment given a natural language description. However, most current systems assume a static environment, which limits the system's applicability in realworld scenarios

Cited by 34SourcecodeScholar
2025

Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments

ICRA 2025

Robot models, particularly those trained with large amounts of data, have recently shown a plethora of real-world manipulation and navigation capabilities. Several independent efforts have shown that given sufficient training data in an environment, robot policies can generalize to demonstrated vari

Cited by 49SourcecodeScholar
2024

OPEN TEACH: A Versatile Teleoperation System for Robotic Manipulation

CoRL 2024poster

Open-sourced, user-friendly tools form the bedrock of scientific advancement across disciplines. The widespread adoption of data-driven learning has led to remarkable progress in multi-fingered dexterity, bimanual manipulation, and applications ranging from logistics to home robotics. However, exist…

Cited by 56SourcecodeScholar
2024

See to Touch: Learning Tactile Dexterity through Visual Incentives

ICRA 2024poster

Equipping multi-fingered robots with tactile sensing is crucial for achieving the precise, contact-rich, and dexterous manipulation that humans excel at. However, relying solely on tactile sensing fails to provide adequate cues for reasoning about objects’ spatial configurations, limiting the abilit…

Cited by 39SourcecodeScholar
2023

CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

RSS 2023poster

We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this…

2023

Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play

CoRL 2023poster

Teaching dexterity to multi-fingered robots has been a longstanding challenge in robotics. Most prominent work in this area focuses on learning controllers or policies that either operate on visual observations or state estimates derived from vision. However, such methods perform poorly on fine-grai…

Cited by 64SourceScholar
2023

Holo-Dex: Teaching Dexterity with Immersive Mixed Reality

ICRA 2023poster

A fundamental challenge in teaching robots is to provide an effective interface for human teachers to demonstrate useful skills to a robot. This challenge is exacerbated in dexterous manipulation, where teaching high-dimensional, contact-rich behaviors often require esoteric teleoperation tools. In…

Cited by 71SourcecodeScholar
2022

Torchaudio: Building Blocks for Audio and Speech Processing

ICASSP 2022accepted

This document describes version 0.10 of TorchAudio: building blocks for machine learning applications in the audio and speech processing domain. The objective of TorchAudio is to accelerate the development and deployment of machine learning applications for researchers and engineers by providing off…

Cited by 0SourceScholar
2021

droidlet: modular, heterogenous, multi-modal agents

ICRA 2021poster

In recent years, there have been significant advances in building end-to-end Machine Learning (ML) systems that learn at scale. But most of these systems are: (a) isolated (perception, speech, or language only); (b) trained on static datasets. On the other hand, in the field of robotics, large-scale…

Cited by 0SourcecodeScholar
2019

PyTorch: An Imperative Style, High-Performance Deep Learning Library

NeurIPS 2019poster

Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it was designed from first principles to support an imperative and Pythonic programming style that supports code as a…

2017

Discovering Causal Signals in Images

CVPR 2017spotlight

This paper establishes the existence of observable footprints that reveal the "causal dispositions" of the object categories appearing in collections of images. We achieve this goal in two steps. First, we take a learning approach to observational causal discovery, and build a classifier that achi…

Cited by 302PDFScholar
2017

Episodic Exploration for Deep Deterministic Policies for StarCraft Micromanagement

ICLR 2017poster

We consider scenarios from the real-time strategy game StarCraft as benchmarks for reinforcement learning algorithms. We focus on micromanagement, that is, the short-term, low-level control of team members during a battle. We propose several scenarios that are challenging for reinforcement learning…

Cited by 21SourceScholar
2015

Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks

NeurIPS 2015poster

In this paper we introduce a generative model capable of producing high quality samples of natural images. Our approach uses a cascade of convolutional networks (convnets) within a Laplacian pyramid framework to generate images in a coarse-to-fine fashion. At each level of the pyramid a separate gen…

Cited by 3162SourcePDFScholar