← Search

Dhruv Batra

113 accepted papers

2024

ASC: Adaptive Skill Coordination for Robotic Mobile Manipulation

RA-L 2024

We present Adaptive Skill Coordination (ASC) – an approach for accomplishing long-horizon tasks like mobile pick-and-place (i.e., navigating to an object, picking it, navigating to another location, and placing it). ASC consists of three components – (1) a library of basic visuomotor <italic xmlns:m

Cited by 74SourceScholar
2024

AutoNeRF: Training Implicit Scene Representations with Autonomous Agents

IROS 2024

Implicit representations such as Neural Radiance Fields (NeRF) allow to map color, density and semantics in a 3D scene through a continuous neural function. However, these models typically require manual and careful human data collection for training. This paper addresses the problem of active explo

Cited by 16SourcecodeScholar
2024

Embodiment Randomization for Cross Embodiment Navigation

IROS 2024poster

We present Embodiment Randomization, a simple, inexpensive, and intuitive technique for training robust behavior policies that can be transferred to multiple robot embodiments. While prior works require real-world data from multiple robots, or complex algorithmic adjustments to address the challenge…

Cited by 0SourceScholar
2024

GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation

CVPR 2024poster

The Embodied AI community has recently made significant strides in visual navigation tasks exploring targets from 3D coordinates objects language description and images. However these navigation models often handle only a single input modality as the target. With the progress achieved so far it is t…

2024

GOAT: GO to Any Thing

RSS 2024poster

In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system cap…

2024

HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation

IROS 2024poster

We present the Habitat-Matterport 3D Open Vocabulary Object Goal Navigation dataset (HM3D-OVON), a large-scale benchmark that broadens the scope and semantic range of prior Object Goal Navigation (ObjectNav) benchmarks. Leveraging the HM3DSem dataset, HM3D-OVON incorporates over 15k annotated instan…

Cited by 10SourceScholar
2024

Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots

ICLR 2024poster

We present Habitat 3.0: a simulation platform for studying collaborative human-robot tasks in home environments. Habitat 3.0 offers contributions across three dimensions: (1) Accurate humanoid simulation: addressing challenges in modeling complex deformable bodies and diversity in appearance and mot…

Cited by 111SourcePDFScholar
2024

Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

CVPR 2024poster

We contribute the Habitat Synthetic Scene Dataset a dataset of 211 high-quality 3D scenes and use it to test navigation agent generalization to realistic 3D environments. Our dataset represents real interiors and contains a diverse set of 18656 models of real-world objects. We investigate the impact…

Cited by 53SourcePDFScholar
2024

IndoorSim-to-OutdoorReal: Learning to Navigate Outdoors Without Any Outdoor Experience

RA-L 2024

We present IndoorSim-to-OutdoorReal (I2O), an end-to-end learned visual navigation approach, trained solely in simulated short-range indoor environments, and demonstrate zero-shot sim-to-real transfer to the outdoors for long-range navigation on the Spot robot. Our method uses zero real-world experi

Cited by 19SourceScholar
2024

OpenEQA: Embodied Question Answering in the Era of Foundation Models

CVPR 2024poster

We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory exemplified by agents on smart glasses or b…

Cited by 118SourcePDFScholar
2024

Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control

NeurIPS 2024spotlight

Embodied AI agents require a fine-grained understanding of the physical world mediated through visual and language inputs. Such capabilities are difficult to learn solely from task-specific data. This has led to the emergence of pre-trained vision-language models as a tool for transferring represent…

2024

Seeing the Unseen: Visual Common Sense for Semantic Placement

CVPR 2024poster

Computer vision tasks typically involve describing what is visible in an image (e.g. classification detection segmentation and captioning). We study a visual common sense task that requires understanding 'what is not visible'. Specifically given an image (e.g. of a living room) and a name of an obje…

Cited by 3SourcePDFScholar
2024

VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation

ICRA 2024poster

Understanding how humans leverage semantic knowledge to navigate unfamiliar environments and decide where to explore next is pivotal for developing robots capable of human-like search behaviors. We introduce a zero-shot navigation approach, Vision-Language Frontier Maps (VLFM), which is inspired by…

Cited by 97SourcecodeScholar
2024

What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?

ICRA 2024poster

We present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluat…

Cited by 6SourceScholar
2023

Adaptive Coordination in Social Embodied Rearrangement

ICML 2023poster

We present the task of "Social Rearrangement", consisting of cooperative everyday tasks like setting up the dinner table, tidying a house or unpacking groceries in a simulated multi-agent environment. In Social Rearrangement, two robots coordinate to complete a long-horizon task, using onboard sensi…

Cited by 12SourcePDFScholar
2023

BC-IRL: Learning Generalizable Reward Functions from Demonstrations

ICLR 2023top-25%

How well do reward functions learned with inverse reinforcement learning (IRL) generalize? We illustrate that state-of-the-art IRL algorithms, which maximize a maximum-entropy objective, learn rewards that overfit to the demonstrations. Such rewards struggle to provide meaningful rewards for states…

Cited by 9SourcePDFScholar
2023

Emergence of Maps in the Memories of Blind Navigation Agents

ICLR 2023top-5%

Animal navigation research posits that organisms build and maintain internal spa- tial representations, or maps, of their environment. We ask if machines – specifically, artificial intelligence (AI) navigation agents – also build implicit (or ‘mental’) maps. A positive answer to this question would…

Cited by 31SourcePDFScholar
2023

FindThis: Language-Driven Object Disambiguation in Indoor Environments

CoRL 2023poster

Natural language is naturally ambiguous. In this work, we consider interactions between a user and a mobile service robot tasked with locating a desired object, specified by a language utterance. We present a task FindThis, which addresses the problem of how to disambiguate and locate the particular…

Cited by 11SourceScholar
2023

Galactic: Scaling End-to-End Reinforcement Learning for Rearrangement at 100k Steps-per-Second

CVPR 2023poster

We present Galactic, a large-scale simulation and reinforcement-learning (RL) framework for robotic mobile manipulation in indoor environments. Specifically, a Fetch robot (equipped with a mobile base, 7DoF arm, RGBD camera, egomotion, and onboard sensing) is spawned in a home environment and asked…

2023

Habitat-Matterport 3D Semantics Dataset

CVPR 2023highlight

We present the Habitat-Matterport 3D Semantics (HM3DSEM) dataset. HM3DSEM is the largest dataset of 3D real-world spaces with densely annotated semantics that is currently available to the academic community. It consists of 142,646 object instance annotations across 216 3D spaces and 3,100 rooms wit…

2023

HomeRobot: Open-Vocabulary Mobile Manipulation

CoRL 2023poster

HomeRobot (noun): An affordable compliant robot that navigates homes and manipulates a wide range of objects in order to complete everyday tasks. Open-Vocabulary Mobile Manipulation (OVMM) is the problem of picking any object in any unseen environment, and placing it in a commanded location. This i…

Cited by 98SourcecodeScholar
2023

Navigating to Objects Specified by Images

ICCV 2023poster

Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform this task in both simulation and the real world. Our modula…

Cited by 41PDFScholar
2023

PIRLNav: Pretraining With Imitation and RL Finetuning for ObjectNav

CVPR 2023poster

We study ObjectGoal Navigation -- where a virtual robot situated in a new environment is asked to navigate to an object. Prior work has shown that imitation learning (IL) using behavior cloning (BC) on a dataset of human demonstrations achieves promising results. However, this has limitations -- 1)…

Cited by 68SourcePDFScholar
2023

Simple and Effective Synthesis of Indoor 3D Scenes

AAAI 2023technical

We study the problem of synthesizing immersive 3D indoor scenes from one or a few images. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the input images while maintaining 3D consistency. Existing approaches are highly…

2023

ViNL: Visual Navigation and Locomotion Over Obstacles

ICRA 2023poster

We present Visual Navigation and Locomotion over obstacles (ViNL), which enables a quadrupedal robot to navigate unseen apartments while stepping over small obstacles that lie in its path (e.g., shoes, toys, cables), similar to how humans and pets lift their feet over objects as they walk. ViNL cons…

Cited by 29SourcecodeScholar
2023

Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?

NeurIPS 2023poster

We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we syste…

Cited by 161SourcePDFScholar
2022

Benchmarking Augmentation Methods for Learning Robust Navigation Agents: the Winning Entry of the 2021 iGibson Challenge

IROS 2022poster

Recent advances in deep reinforcement learning and scalable photorealistic simulation have led to increasingly mature embodied AI for various visual tasks, including navigation. However, while impressive progress has been made for teaching embodied agents to navigate static environments, much less p…

Cited by 11SourceScholar
2022

Cross-Domain Transfer via Semantic Skill Imitation

CoRL 2022poster

We propose an approach for semantic imitation, which uses demonstrations from a source domain, e.g. human videos, to accelerate reinforcement learning (RL) in a different target domain, e.g. a robotic manipulator in a simulated kitchen. Instead of imitating low-level actions like joint velocities, o…

Cited by 19SourceScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Episodic Memory Question Answering

CVPR 2022oral

Egocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human communicates with an AI agent powering such a device by asking questions (e.g., "where did you last see my keys?"). In orde…

Cited by 40PDFScholar
2022

Habitat-Web: Learning Embodied Object-Search Strategies From Human Demonstrations at Scale

CVPR 2022poster

We present a large-scale study of imitating human demonstrations on tasks that require a virtual robot to search for objects in new environments - (1) ObjectGoal Navigation (e.g. 'find & go to a chair') and (2) Pick&Place (e.g. 'find mug, pick mug, find counter, place mug on counter'). First, we dev…

Cited by 117PDFcodeScholar
2022

Housekeep: Tidying Virtual Households Using Commonsense Reasoning

ECCV 2022poster

"We introduce Housekeep, a benchmark to evaluate commonsense reasoning in the home for embodied AI. In Housekeep, an embodied agent must tidy a house by rearranging misplaced objects without explicit instructions specifying which objects need to be rearranged. Instead, the agent must learn from and…

2022

Is Mapping Necessary for Realistic PointGoal Navigation?

CVPR 2022poster

Can an autonomous agent navigate in a new environment without building an explicit map? For the task of PointGoal navigation ('Go to (x, y)') under idealized settings (no RGB-D and actuation noise, perfect GPS+Compass), the answer is a clear 'yes' - map-less neural models composed of task-agnostic c…

Cited by 54PDFcodeScholar
2022

Memory-Augmented Reinforcement Learning for Image-Goal Navigation

IROS 2022poster

In this work, we present a memory-augmented approach for image-goal navigation. Earlier attempts, including RL-based and SLAM-based approaches have either shown poor generalization performance, or are heavily-reliant on pose/depth sensors. Our method is based on an attention-based end-to-end model t…

Cited by 88SourcecodeScholar
2022

Rethinking Sim2Real: Lower Fidelity Simulation Leads to Higher Sim2Real Transfer in Navigation

CoRL 2022poster

If we want to train robots in simulation before deploying them in reality, it seems natural and almost self-evident to presume that reducing the sim2real gap involves creating simulators of increasing fidelity (since reality is what it is). We challenge this assumption and present a contrary hypothe…

Cited by 47SourceScholar
2022

SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

NeurIPS 2022accept

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D vi…

Cited by 94SourcePDFScholar
2022

VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied Rearrangement

NeurIPS 2022accept

We present Variable Experience Rollout (VER), a technique for efficiently scaling batched on-policy reinforcement learning in heterogenous environments (where different environments take vastly different times to generate rollouts) to many GPUs residing on, potentially, many machines. VER combines t…

2022

ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

NeurIPS 2022accept

We present a scalable approach for learning open-world object-goal navigation (ObjectNav) – the task of asking a virtual robot (agent) to find any instance of an object in an unexplored environment (e.g., “find a sink”). Our approach is entirely zero-shot – i.e., it does not require ObjectNav reward…

2021

Contrast and Classify: Training Robust VQA Models

ICCV 2021poster

Recent Visual Question Answering (VQA) models have shown impressive performance on the VQA benchmark but remain sensitive to small linguistic variations in input questions. Existing approaches address this by augmenting the dataset with question paraphrases from visual question generation models or…

Cited by 35PDFcodeScholar
2021

Habitat 2.0: Training Home Assistants to Rearrange their Habitat

NeurIPS 2021spotlight

We introduce Habitat 2.0 (H2.0), a simulation platform for training virtual robots in interactive 3D environments and complex physics-enabled scenarios. We make comprehensive contributions to all levels of the embodied AI stack – data, simulation, and benchmark tasks. Specifically, we present: (i) R…

2021

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

NeurIPS 2021poster

We present the Habitat-Matterport 3D (HM3D) dataset. HM3D is a large-scale dataset of 1,000 building-scale 3D reconstructions from a diverse set of real-world locations. Each scene in the dataset consists of a textured 3D mesh reconstruction of interiors such as multi-floor residences, stores, and ot…

Cited by 434SourcecodeScholar
2021

Large Batch Simulation for Deep Reinforcement Learning

ICLR 2021poster

We accelerate deep reinforcement learning-based training in visually complex 3D environments by two orders of magnitude over prior work, realizing end-to-end training speeds of over 19,000 frames of experience per second on a single GPU and up to 72,000 frames per second on a single eight-GPU machin…

2021

Learning Navigation Skills for Legged Robots with Learned Robot Embeddings

IROS 2021poster

Recent work has shown results on learning navigation policies for idealized cylinder agents in simulation and transferring them to real wheeled robots. Deploying such navigation policies on legged robots can be challenging due to their complex dynamics, and the large dynamical difference between cyl…

Cited by 21SourceScholar
2021

SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation

NeurIPS 2021poster

Natural language instructions for visual navigation often use scene descriptions (e.g., bedroom) and object references (e.g., green chairs) to provide a breadcrumb trail to a goal location. This work presents a transformer-based vision-and-language navigation (VLN) agent that uses two different visu…

Cited by 64SourcePDFScholar
2021

SOrT-ing VQA Models : Contrastive Gradient Learning for Improved Consistency

NAACL 2021long

Recent research in Visual Question Answering (VQA) has revealed state-of-the-art models to be inconsistent in their understanding of the world - they answer seemingly difficult questions requiring reasoning correctly but get simpler associated sub-questions wrong. These sub-questions pertain to lowe…

2021

Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

AAAI 2021technical

We study the task of semantic mapping – specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (‘what is where?’) from egocentric observations of an RGB-D camera with known pose (via localiz…

2021

Success Weighted by Completion Time: A Dynamics-Aware Evaluation Criteria for Embodied Navigation

IROS 2021poster

We present Success weighted by Completion Time (SCT), a new metric for evaluating navigation performance for mobile robots. Several related works on navigation have used Success weighted by Path Length (SPL) as the primary method of evaluating the path an agent makes to a goal location, but SPL is l…

Cited by 27SourceScholar
2021

THDA: Treasure Hunt Data Augmentation for Semantic Navigation

ICCV 2021poster

Can general-purpose neural models learn to navigate? For PointGoal navigation (""go to x, y""), the answer is a clear `yes' -- mapless neural models composed of task-agnostic components (CNNs and RNNs) trained with large-scale model-free reinforcement learning achieve near-perfect performance. Howev…

Cited by 91PDFScholar
2021

The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal Navigation

ICCV 2021poster

It is fundamental for personal robots to reliably navigate to a specified goal. To study this task, PointGoal navigation has been introduced in simulated Embodied AI environments. Recent advances solve this PointGoal navigation task with near-perfect accuracy (99.6% success) in photo-realistically s…

Cited by 53PDFScholar
2021

Waypoint Models for Instruction-Guided Navigation in Continuous Environments

ICCV 2021poster

Little inquiry has explicitly addressed the role of action spaces in language-guided visual navigation -- either in terms of its effect on navigation success or the efficiency with which a robotic agent could execute the resulting trajectory. Building on the recently released VLN-CE setting for inst…

Cited by 94PDFcodeScholar
2020

Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments

ECCV 2020poster

We develop a language-guided navigation task set in a continuous 3D environment where agents must execute low-level actions to follow natural language navigation directions. By being situated in continuous environments, this setting lifts a number of assumptions implicit in prior work that represent…

2020

DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

ICLR 2020poster

We present Decentralized Distributed Proximal Policy Optimization (DD-PPO), a method for distributed reinforcement learning in resource-intensive simulated environments. DD-PPO is distributed (uses multiple machines), decentralized (lacks a centralized server), and synchronous (no computation is eve…

Cited by 542SourcecodeScholar
2020

Dialog without Dialog Data: Learning Visual Dialog Agents from VQA Data

NeurIPS 2020poster

Can we develop visually grounded dialog agents that can efficiently adapt to new tasks without forgetting how to talk to people? Such agents could leverage a larger variety of existing data to generalize to a new task, minimizing expensive data collection and annotation. In this work, we study a set…

2020

Embodied Multimodal Multitask Learning

IJCAI 2020poster

Visually-grounded embodied language learning models have recently shown to be effective at learning multiple multimodal tasks such as following navigational instructions and answering questions. In this paper, we address two key limitations of these models, (a) the inability to transfer the grounded…

Cited by 0SourcePDFScholar
2020

IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL

IJCAI 2020poster

We propose a novel framework to identify sub-goals useful for exploration in sequential decision making tasks under partial observability. We utilize the variational intrinsic control framework (Gregor et.al., 2016) which maximizes empowerment -- the ability to reliably reach a diverse set of states…

2020

Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

ECCV 2020poster

Following a navigation instruction such as 'Walk down the stairs and stop at the brown sofa' requires embodied AI agents to ground referenced scene elements referenced (e.g. 'stairs') to visual content in the environment (pixels corresponding to 'stairs'). We ask the following question -- can we lev…

2020

Integrating Egocentric Localization for More Realistic Point-Goal Navigation Agents

CoRL 2020

Recent work has presented embodied agents that can navigate to point-goal targets in novel indoor environments with near-perfect accuracy. However, these agents are equipped with idealized sensors for localization and take deterministic actions. This setting is practically sterile by comparison to t

Cited by 0SourcePDFScholar
2020

Large-scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline

ECCV 2020poster

Prior work in visual dialog has focused on training deep neural models on VisDial in isolation. Instead, we present an approach to leverage pretraining on related vision-language datasets before transferring to visual dialog. We adapt the recently proposed ViLBERT model (Lu et al. 2019) for multi-tu…

2020

Seeing the Un-Scene: Learning Amodal Semantic Maps for Room Navigation

ECCV 2020poster

We introduce a learning-based approach for room navigation using semantic maps. Our proposed architecture learns to predict top-down belief maps of regions that lie beyond the agent’s field of view while modeling architectural and stylistic regularities in houses. First, we train a model to generate…

Cited by 70SourcePDFScholar
2020

Sim-to-Real Transfer for Vision-and-Language Navigation

CoRL 2020

We study the challenging problem of releasing a robot in a previously unseen environment, and having it follow unconstrained natural language navigation instructions. Recent work on the task of Vision-and-Language Navigation (VLN) has achieved significant progress in simulation. To assess the implic

2020

Sim2Real Predictivity: Does Evaluation in Simulation Predict Real-World Performance?

RA-L 2020

Does progress in simulation translate to progress on robots? If one method outperforms another in simulation, how likely is that trend to hold in reality on a robot? We examine this question for embodied PointGoal navigation - developing engineering tools and a research paradigm for evaluating a sim

Cited by 252SourceScholar
2020

Spatially Aware Multimodal Transformers for TextVQA

ECCV 2020poster

Textual cues are essential for everyday tasks like buying groceries and using public transport. To develop this assistive technology, we study the TextVQA task, i.e., reasoning about text in images to answer a question. Existing approaches are limited in their use of spatial relations and rely on fu…

2019

Audio Visual Scene-Aware Dialog

CVPR 2019poster

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video whi…

Cited by 226PDFcodeScholar
2019

Chasing Ghosts: Instruction Following as Bayesian State Tracking

NeurIPS 2019poster

A visually-grounded navigation instruction can be interpreted as a sequence of expected observations and actions an agent following the correct trajectory would encounter and perform. Based on this intuition, we formulate the problem of finding the goal location in Vision-and-Language Navigation (VL…

2019

Embodied Amodal Recognition: Learning to Move to Perceive Objects

ICCV 2019poster

Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively control the viewing angle to better understand object shapes and semantics. In this…

Cited by 75PDFScholar
2019

Embodied Question Answering in Photorealistic Environments With Point Cloud Perception

CVPR 2019oral

To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments (Matterport 3D). We thoroughly study navigation policies that utilize 3D poin…

Cited by 193PDFScholar
2019

End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features

ICASSP 2019accepted

In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differenti…

Cited by 0SourceScholar
2019

Habitat: A Platform for Embodied AI Research

ICCV 2019oral

We present Habitat, a platform for research in embodied artificial intelligence (AI). Habitat enables training embodied agents (virtual robots) in highly efficient photorealistic 3D simulation. Specifically, Habitat consists of: (i) Habitat-Sim: a flexible, high-performance 3D simulator with configu…

Cited by 2011PDFcodeScholar
2019

Modeling the Long Term Future in Model-Based Reinforcement Learning

ICLR 2019poster

In model-based reinforcement learning, the agent interleaves between model learning and planning. These two components are inextricably intertwined. If the model is not able to provide sensible long-term prediction, the executed planer would exploit model flaws, which can yield catastrophic failur…

Cited by 42SourcePDFScholar
2019

Multi-Target Embodied Question Answering

CVPR 2019poster

Embodied Question Answering (EQA) is a relatively new task where an agent is asked to answer questions about its environment from egocentric perception. EQA as introduced in [8] makes the fundamental assumption that every question, e.g., "what color is the car?", has exactly one target ("car") bein…

Cited by 130PDFcodeScholar
2019

Probabilistic Neural Symbolic Models for Interpretable Visual Question Answering

ICML 2019oral

We propose a new class of probabilistic neural-symbolic models, that have symbolic functional programs as a latent, stochastic variable. Instantiated in the context of visual question answering, our probabilistic formulation offers two key conceptual advantages over prior neural-symbolic models for…

Cited by 109SourcePDFScholar
2019

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

ICCV 2019poster

Diverse and accurate vision+language modeling is an important goal to retain creative freedom and maintain user engagement. However, adequately capturing the intricacies of diversity in language models is challenging. Recent works commonly resort to latent variable models augmented with more or less…

Cited by 84PDFScholar
2019

SplitNet: Sim2Sim and Task2Task Transfer for Embodied Visual Navigation

ICCV 2019poster

We propose SplitNet, a method for decoupling visual perception and policy learning. By incorporating auxiliary tasks and selective learning of portions of the model, we explicitly decompose the learning objectives for visual navigation into perceiving the world and acting on that perception. We show…

Cited by 78PDFcodeScholar
2019

Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded

ICCV 2019poster

Many vision and language models suffer from poor visual grounding -- often falling back on easy-to-learn language priors rather than basing their decisions on visual concepts in the image. In this work, we propose a generic approach called Human Importance-aware Network Tuning (HINT) that effectivel…

Cited by 305PDFScholar
2019

Towards VQA Models That Can Read

CVPR 2019poster

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to…

Cited by 1328PDFcodeScholar
2019

Trainable Decoding of Sets of Sequences for Neural Sequence Models

ICML 2019oral

Many sequence prediction tasks admit multiple correct outputs and so, it is often useful to decode a set of outputs that maximize some task-specific set-level metric. However, retooling standard sequence prediction procedures tailored towards predicting the single best output leads to the decoding o…

Cited by 3SourcePDFScholar
2019

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

NeurIPS 2019poster

We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, processing both visual and textual inputs in separate streams that inter…

Cited by 4490SourcePDFScholar
2019

nocaps: novel object captioning at scale

ICCV 2019poster

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supe…

Cited by 420PDFcodeScholar
2018

Choose Your Neuron: Incorporating Domain Knowledge through Neuron-Importance

ECCV 2018poster

Individual neurons in convolutional neural networks supervised for image-level classification tasks have been shown to implicitly learn semantically meaningful concepts ranging from simple textures and shapes to whole or partial objects – forming a “dictionary” of concepts acquired through the learn…

2018

Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

CVPR 2018poster

A number of studies have found that today's Visual Question Answering (VQA) models are heavily driven by superficial correlations in the training data and lack sufficient image grounding. To encourage development of models geared towards the latter, we propose a new setting for VQA where for every q…

Cited by 772SourcePDFScholar
2018

Learn from Your Neighbor: Learning Multi-modal Mappings from Sparse Annotations

ICML 2018oral

Many structured prediction problems (particularly in vision and language domains) are ambiguous, with multiple outputs being ‘correct’ for an input {–} e.g. there are many ways of describing an image, multiple ways of translating a sentence; however, exhaustively annotating the applicability of all…

Cited by 6SourcePDFScholar
2018

Neural Modular Control for Embodied Question Answering

CoRL 2018

We present a modular approach for learning policies for navigation over long planning horizons from language input. Our hierarchical policy operates at multiple timescales, where the higher-level master policy proposes subgoals to be executed by specialized sub-policies. Our choice of subgoals is co

2018

Neural-Guided Deductive Search for Real-Time Program Synthesis from Examples

ICLR 2018poster

Synthesizing user-intended programs from a small number of input-output exam- ples is a challenging problem with several important applications like spreadsheet manipulation, data wrangling and code refactoring. Existing synthesis systems either completely rely on deductive logic techniques that are…

Cited by 190SourcePDFScholar
2018

Visual Coreference Resolution in Visual Dialog using Neural Module Networks

ECCV 2018poster

Visual dialog entails answering a series of questions grounded in an image, using dialog history as context. In addition to the challenges found in visual question answering (VQA), which can be seen as one-round dialog, visual dialog encompasses several more. We focus on one such problem called ‘vis…

2018

Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition

CoRL 2018

In an open-world setting, it is inevitable that an intelligent agent (e.g., a robot) will encounter visual objects, attributes or relationships it does not recognize. In this work, we develop an agent empowered with visual curiosity, i.e. the ability to ask questions to an Oracle (e.g., human) about

2017

Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model

NeurIPS 2017poster

We present a novel training framework for neural sequence models, particularly for grounded dialog generation. The standard training paradigm for these models is maximum likelihood estimation (MLE), or minimizing the cross-entropy of the human responses. Across a variety of domains, a recurring prob…

2017

Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-In-The-Blank Image Captioning

CVPR 2017poster

We develop the first approximate inference algorithm for 1-Best (and M-Best) decoding in bidirectional neural sequence models by extending Beam Search (BS) to reason about both forward and backward time dependencies. Beam Search (BS) is a widely used approximate inference algorithm for decoding se…

Cited by 56PDFScholar
2017

Counting Everyday Objects in Everyday Scenes

CVPR 2017spotlight

We are interested in counting the number of instances of object classes in natural, everyday images. Previous counting approaches tackle the problem in restricted domains such as counting pedestrians in surveillance videos. Counts can also be estimated from outputs of other vision tasks like object…

Cited by 205PDFcodeScholar
2017

Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization

ICCV 2017poster

We propose a technique for producing 'visual explanations' for decisions from a large class of Convolutional Neural Network (CNN)-based models, making them more transparent. Our approach - Gradient-weighted Class Activation Mapping (Grad-CAM), uses the gradients of any target concept (say logits for…

Cited by 24144PDFcodeScholar
2017

LR-GAN: Layered Recursive Generative Adversarial Networks for Image Generation

ICLR 2017poster

We present LR-GAN: an adversarial image generation model which takes scene structure and context into account. Unlike previous generative adversarial networks (GANs), the proposed GAN learns to generate image background and foregrounds separately and recursively, and stitch the foregrounds on the ba…

Cited by 297SourcecodeScholar
2017

Learning Cooperative Visual Dialog Agents With Deep Reinforcement Learning

ICCV 2017oral

We introduce the first goal-driven training for visual question answering and dialog agents. Specifically, we pose a cooperative `image guessing' game between two agents -- Qbot and Abot -- who communicate in natural language dialog so that Qbot can select an unseen image from a lineup of images. We…

Cited by 493PDFcodeScholar
2017

Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

CVPR 2017poster

Problems at the intersection of vision and language are of significant importance both as challenging research questions and for the rich set of applications they enable. However, inherent structure in our world and bias in our language tend to be a simpler signal for learning than visual modalities…

Cited by 3740PDFScholar
2016

Hierarchical Question-Image Co-Attention for Visual Question Answering

NeurIPS 2016poster

A number of recent works have proposed attention models for Visual Question Answering (VQA) that generate spatial maps highlighting image regions relevant to answering the question. In this paper, we argue that in addition to modeling "where to look" or visual attention, it is equally important to m…

2016

Stochastic Multiple Choice Learning for Training Diverse Deep Ensembles

NeurIPS 2016poster

Many practical perception systems exist within larger processes which often include interactions with users or additional components that are capable of evaluating the quality of predicted solutions. In these contexts, it is beneficial to provide these oracle mechanisms with multiple highly likely h…

Cited by 233SourcePDFScholar
2016

We Are Humor Beings: Understanding and Predicting Visual Humor

CVPR 2016spotlight

Humor is an integral part of human lives. Despite being tremendously impactful, it is perhaps surprising that we do not have a detailed understanding of humor yet. As interactions between humans and AI systems increase, it is imperative that these systems are taught to understand subtleties of human…

Cited by 69PDFcodeScholar
2016

Yin and Yang: Balancing and Answering Binary Visual Questions

CVPR 2016poster

The complex compositional structure of language makes problems at the intersection of vision and language challenging. But language also provides a strong prior that can result in good superficial performance, without the underlying models truly understanding the visual content. This can hinder pro…

Cited by 439PDFScholar
2015

Active Learning for Structured Probabilistic Models With Histogram Approximation

CVPR 2015poster

Abstract. This paper studies active learning in structured probabilistic models such as Conditional Random Fields (CRFs). This is a challenging problem because unlike unstructured prediction problems such as binary or multi-class classification, structured prediction problems involve a distribution…

Cited by 32SourcePDFScholar
2015

VQA: Visual Question Answering

ICCV 2015poster

We propose the task of free-form and open-ended Visual Question Answering (VQA). Given an image and a natural language question about the image, the task is to provide an accurate natural language answer. Mirroring real-world scenarios, such as helping the visually impaired, both the questions and a…

Cited by 7071PDFcodeScholar