← Search

Ashish Kapoor

40 accepted papers

2026

DreamControl: Human-Inspired Whole-Body Humanoid Control for Scene Interaction Via Guided Diffusion

ICRA 2026poster

We introduce DreamControl, a novel methodology for learning autonomous whole-body humanoid skills. DreamControl leverages the strengths of diffusion models and Reinforcement Learning (RL): our core innovation is the use of a diffusion prior trained on human motion data, which subsequently guides an …

2024

ConBaT: Control Barrier Transformer for Safe Robot Learning from Demonstrations

ICRA 2024poster

Large-scale self-supervised models have recently revolutionized our ability to perform a variety of tasks within the vision and language domains. However, using such models for autonomous systems is challenging because of safety requirements: besides executing correct actions, an autonomous agent mu…

Cited by 1SourceScholar
2023

ClimaX: A foundation model for weather and climate

ICML 2023poster

Recent data-driven approaches based on machine learning aim to directly solve a downstream forecasting or projection task by learning a data-driven functional mapping using deep neural networks. However, these networks are trained using curated and homogeneous climate datasets for specific spatiotem…

2023

Is Imitation All You Need? Generalized Decision-Making with Dual-Phase Training

ICCV 2023poster

We introduce DualMind, a generalist agent designed to tackle various decision-making tasks that addresses challenges posed by current methods, such as overfitting behaviors and dependence on task-specific fine-tuning. DualMind uses a novel "Dual-phase" training strategy that emulates how humans lear…

Cited by 17PDFcodeScholar
2023

LATTE: LAnguage Trajectory TransformEr

ICRA 2023poster

Natural language is one of the most intuitive ways to express human intent. However, translating instructions and commands towards robotic motion generation and deployment in the real world is far from being an easy task. The challenge of combining a robot's inherent low-level geometric and kinodyna…

Cited by 79SourcecodeScholar
2023

PACT: Perception-Action Causal Transformer for Autoregressive Robotics Pre-Training

IROS 2023poster

Robotics has long been a field riddled with complex systems architectures whose modules and connections, whether traditional or learning-based, require significant human expertise and prior knowledge. Inspired by large pre-trained language models, this work introduces a paradigm for pretraining a ge…

Cited by 20SourceScholar
2023

SMART: Self-supervised Multi-task pretrAining with contRol Transformers

ICLR 2023top-25%

Self-supervised pretraining has been extensively studied in language and vision domains, where a unified model can be easily adapted to various downstream tasks by pretraining representations without explicit labels. When it comes to sequential decision-making tasks, however, it is difficult to prop…

2022

3DB: A Framework for Debugging Computer Vision Models

NeurIPS 2022accept

We introduce 3DB: an extendable, unified framework for testing and debugging vision models using photorealistic simulation. We demonstrate, through a wide range of use cases, that 3DB allows users to discover vulnerabilities in computer vision systems and gain insights into how models make decision…

2022

COMPASS: Contrastive Multimodal Pretraining for Autonomous Systems

IROS 2022poster

Learning representations that generalize across tasks and domains is challenging yet necessary for autonomous systems. Although task-driven approaches are appealing, de-signing models specific to each application can be difficult in the face of limited data, especially when dealing with highly varia…

Cited by 10SourcecodeScholar
2022

Learning to Simulate Realistic LiDARs

IROS 2022poster

Simulating realistic sensors is a challenging part in data generation for autonomous systems, often involving carefully handcrafted sensor design, scene properties, and physics modeling. To alleviate this, we introduce a pipeline for data-driven simulation of a realistic LiDAR sensor. We propose a m…

Cited by 19SourceScholar
2022

Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using Transformers

IROS 2022poster

Natural language is the most intuitive medium for us to interact with other people when expressing commands and instructions. However, using language is seldom an easy task when humans need to express their intent towards robots, since most of the current language interfaces require rigid templates…

Cited by 60SourcecodeScholar
2021

Adversarial Training on Point Clouds for Sim-to-Real 3D Object Detection

RA-L 2021

In this work we address the problem of 3D object detection from point clouds in data-limited environments. Training with simulated data is a common approach in such scenarios; however a sim-to-real gap exists between clean and crisp simulated clouds and noisy real clouds. Previous sim-to-real approa

Cited by 22SourceScholar
2021

Quantum algorithms for reinforcement learning with a generative model

ICML 2021spotlight

Reinforcement learning studies how an agent should interact with an environment to maximize its cumulative reward. A standard way to study this question abstractly is to ask how many samples an agent needs from the environment to learn an optimal policy for a $\gamma$-discounted Markov decision proc…

Cited by 38SourcePDFScholar
2021

Unadversarial Examples: Designing Objects for Robust Vision

NeurIPS 2021poster

We study a class of computer vision settings wherein one can modify the design of the objects being recognized. We develop a framework that leverages this capability---and deep networks' unusual sensitivity to input perturbations---to design ``robust objects,'' i.e., objects that are explicitly opti…

Cited by 58SourcePDFScholar
2020

Denoised Smoothing: A Provable Defense for Pretrained Classifiers

NeurIPS 2020poster

We present a method for provably defending any pretrained image classifier against $\ell_p$ adversarial attacks. This method, for instance, allows public vision API providers and users to seamlessly convert pretrained non-robust classification services into provably robust ones. By prepending a cust…

2020

Do Adversarially Robust ImageNet Models Transfer Better?

NeurIPS 2020oral

Transfer learning is a widely-used paradigm in deep learning, where models pre-trained on standard datasets can be efficiently adapted to downstream tasks. Typically, better pre-trained models yield better transfer results, suggesting that initial accuracy is a key aspect of transfer learning perfor…

2020

Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal Representations

IROS 2020poster

Machines are a long way from robustly solving open-world perception-control tasks, such as first-person view (FPV) aerial navigation. While recent advances in end-to- end Machine Learning, especially Imitation Learning and Reinforcement appear promising, they are constrained by the need of large amo…

Cited by 61SourcecodeScholar
2020

Multi-Robot Collision Avoidance under Uncertainty with Probabilistic Safety Barrier Certificates

NeurIPS 2020spotlight

Safety in terms of collision avoidance for multi-robot systems is a difficult challenge under uncertainty, non-determinism, and lack of complete information. This paper aims to propose a collision avoidance method that accounts for both measurement uncertainty and motion uncertainty. In particular,…

2020

Safety Considerations in Deep Control Policies with Safety Barrier Certificates Under Uncertainty

IROS 2020poster

Recent advances in Deep Machine Learning have shown promise in solving complex perception and control loops via methods such as reinforcement and imitation learning. However, guaranteeing safety for such learned deep policies has been a challenge due to issues such as partial observability and diffi…

Cited by 5SourceScholar
2020

TartanAir: A Dataset to Push the Limits of Visual SLAM

IROS 2020poster

We present a challenging dataset, the TartanAir, for robot navigation tasks and more. The data is collected in photo-realistic simulation environments with the presence of moving objects, changing light and various weather conditions. By collecting data in simulations, we are able to obtain multi-mo…

Cited by 406SourcecodeScholar
2019

Bias Correction of Learned Generative Models using Likelihood-Free Importance Weighting

NeurIPS 2019poster

A learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio under model and true distributions. When the likelihood ratio…

Cited by 156SourcePDFScholar
2019

Characterizing Bias in Classifiers using Generative Models

NeurIPS 2019poster

Models that are learned from real-world data are often biased because the data used to train them is biased. This can propagate systemic human biases that exist and ultimately lead to inequitable treatment of people, especially minorities. To characterize bias in learned classifiers, existing approa…

2019

Inverse Optimal Planning for Air Traffic Control

IROS 2019poster

We envision a system that concisely describes the rules of air traffic control, assists human operators and supports dense autonomous air traffic around commercial airports. We develop a method to learn the rules of air traffic control from real data as a cost function via maximum entropy inverse re…

Cited by 8SourcecodeScholar
2019

Visceral Machines: Risk-Aversion in Reinforcement Learning with Intrinsic Physiological Rewards

ICLR 2019poster

As people learn to navigate the world, autonomic nervous system (e.g., ``fight or flight) responses provide intrinsic feedback about the potential consequence of action choices (e.g., becoming nervous when close to a cliff edge or driving fast around a bend.) Physiological changes are correlated wit…

Cited by 19SourcePDFScholar
2018

Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility

ECCV 2018poster

Camera equipped drones are nowadays being used to explore large scenes and reconstruct detailed 3D maps. When free space in the scene is approximately known, an offline planner can generate optimal plans to efficiently explore the scene. However, for exploring unknown scenes, the planner must predic…

Cited by 62SourcePDFScholar
2018

Verifying Controllers Against Adversarial Examples with Bayesian Optimization

ICRA 2018poster

Recent successes in reinforcement learning have lead to the development of complex controllers for realworld robots. As these robots are deployed in safety-critical applications and interact with humans, it becomes critical to ensure safety in order to avoid causing harm. A first step in this direct…

Cited by 63SourcecodeScholar
2017

Adaptive Information Gathering via Imitation Learning

RSS 2017poster

In the adaptive information gathering problem, a policy is required to select an informative sensing location using the history of measurements acquired thus far. While there is an extensive amount of prior work investigating effective practical approximations using variants of Shannon's entropy, th…

Cited by 25SourcePDFScholar
2017

No-regret replanning under uncertainty

ICRA 2017poster

This paper explores the problem of path planning under uncertainty. Specifically, we consider online receding horizon based planners that need to operate in a latent environment where the latent information can be modelled via Gaussian Processes. Online path planning in latent environments is challe…

Cited by 14SourceScholar
2017

Submodular Trajectory Optimization for Aerial 3D Scanning

ICCV 2017poster

Drones equipped with cameras are emerging as a powerful tool for large-scale aerial 3D scanning, but existing automatic flight planners do not exploit all available information about the scene, and can therefore produce inaccurate and incomplete 3D models. We present an automatic method to generate…

Cited by 177PDFScholar