← Search

Abhishek Das

21 accepted papers

2025

UMA: A Family of Universal Models for Atoms

NeurIPS 2025spotlight

The ability to quickly and accurately compute properties from atomic simulations is critical for advancing a large number of applications in chemistry and materials science including drug discovery, energy storage, and semiconductor manufacturing. To address this need, we present a family of Univers…

Cited by 0SourceScholar
2024

EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations

ICLR 2024poster

Equivariant Transformers such as Equiformer have demonstrated the efficacy of applying Transformers to the domain of 3D atomistic systems. However, they are limited to small degrees of equivariant representations due to their computational complexity. In this paper, we investigate whether these arch…

2024

HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation

IROS 2024poster

We present the Habitat-Matterport 3D Open Vocabulary Object Goal Navigation dataset (HM3D-OVON), a large-scale benchmark that broadens the scope and semantic range of prior Object Goal Navigation (ObjectNav) benchmarks. Leveraging the HM3DSem dataset, HM3D-OVON incorporates over 15k annotated instan…

Cited by 10SourceScholar
2023

PIRLNav: Pretraining With Imitation and RL Finetuning for ObjectNav

CVPR 2023poster

We study ObjectGoal Navigation -- where a virtual robot situated in a new environment is asked to navigate to an object. Prior work has shown that imitation learning (IL) using behavior cloning (BC) on a dataset of human demonstrations achieves promising results. However, this has limitations -- 1)…

Cited by 68SourcePDFScholar
2022

Habitat-Web: Learning Embodied Object-Search Strategies From Human Demonstrations at Scale

CVPR 2022poster

We present a large-scale study of imitating human demonstrations on tasks that require a virtual robot to search for objects in new environments - (1) ObjectGoal Navigation (e.g. 'find & go to a chair') and (2) Pick&Place (e.g. 'find mug, pick mug, find counter, place mug on counter'). First, we dev…

Cited by 117PDFcodeScholar
2022

Spherical Channels for Modeling Atomic Interactions

NeurIPS 2022accept

Modeling the energy and forces of atomic systems is a fundamental problem in computational chemistry with the potential to help address many of the world’s most pressing problems, including those related to energy scarcity and climate change. These calculations are traditionally performed using Dens…

2022

Towards Training Billion Parameter Graph Neural Networks for Atomic Simulations

ICLR 2022poster

Recent progress in Graph Neural Networks (GNNs) for modeling atomic simulations has the potential to revolutionize catalyst discovery, which is a key step in making progress towards the energy breakthroughs needed to combat climate change. However, the GNNs that have proven most effective for this t…

Cited by 33SourcePDFScholar
2020

IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL

IJCAI 2020poster

We propose a novel framework to identify sub-goals useful for exploration in sequential decision making tasks under partial observability. We utilize the variational intrinsic control framework (Gregor et.al., 2016) which maximizes empowerment -- the ability to reliably reach a diverse set of states…

2020

Large-scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline

ECCV 2020poster

Prior work in visual dialog has focused on training deep neural models on VisDial in isolation. Instead, we present an approach to leverage pretraining on related vision-language datasets before transferring to visual dialog. We adapt the recently proposed ViLBERT model (Lu et al. 2019) for multi-tu…

2020

Probing Emergent Semantics in Predictive Agents via Question Answering

ICML 2020poster

Recent work has shown how predictive modeling can endow agents with rich knowledge of their surroundings, improving their ability to act in complex environments. We propose question-answering as a general paradigm to decode and understand the representations that such agents develop, applying our me…

Cited by 22SourcePDFScholar
2019

Audio Visual Scene-Aware Dialog

CVPR 2019poster

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video whi…

Cited by 226PDFcodeScholar
2019

Embodied Question Answering in Photorealistic Environments With Point Cloud Perception

CVPR 2019oral

To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments (Matterport 3D). We thoroughly study navigation policies that utilize 3D poin…

Cited by 193PDFScholar
2019

End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features

ICASSP 2019accepted

In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differenti…

Cited by 0SourceScholar
2018

Neural Modular Control for Embodied Question Answering

CoRL 2018

We present a modular approach for learning policies for navigation over long planning horizons from language input. Our hierarchical policy operates at multiple timescales, where the higher-level master policy proposes subgoals to be executed by specialized sub-policies. Our choice of subgoals is co

2017

Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization

ICCV 2017poster

We propose a technique for producing 'visual explanations' for decisions from a large class of Convolutional Neural Network (CNN)-based models, making them more transparent. Our approach - Gradient-weighted Class Activation Mapping (Grad-CAM), uses the gradients of any target concept (say logits for…

Cited by 24144PDFcodeScholar
2017

Learning Cooperative Visual Dialog Agents With Deep Reinforcement Learning

ICCV 2017oral

We introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative `image guessing' game between two agents -- Qbot and Abot -- who communicate in natural language dialog so that Qbot can select an unseen image from a lineup of images. We…

Cited by 493PDFcodeScholar