← Search

Govind Thattai

8 accepted papers

2023

Alexa Arena: A User-Centric Interactive Platform for Embodied AI

NeurIPS 2023poster

We introduce Alexa Arena, a user-centric simulation platform to facilitate research in building assistive conversational embodied agents. Alexa Arena features multi-room layouts and an abundance of interactable objects. With user-friendly graphics and control mechanisms, the platform supports the de…

2023

GIVL: Improving Geographical Inclusivity of Vision-Language Models With Pre-Training Methods

CVPR 2023poster

A key goal for the advancement of AI is to develop technologies that serve the needs not just of one group but of all communities regardless of their geographical region. In fact, a significant proportion of knowledge is locally shared by people from certain regions but may not apply equally in othe…

2023

LEMMA: Learning Language-Conditioned Multi-Robot Manipulation

RA-L 2023

Complex manipulation tasks often require robots with complementary capabilities to collaborate. We introduce a benchmark for <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">L</u> anguag <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xml

Cited by 15SourceScholar
2023

Neural Architecture Search for Parameter-Efficient Fine-tuning of Large Pre-trained Language Models

ACL 2023findings

Parameter-efficient tuning (PET) methods fit pre-trained language models (PLMs) to downstream tasks by either computing a small compressed update for a subset of model parameters, or appending and fine-tuning a small number of new model parameters to the pre-trained network. Hand-designed PET archit…

Cited by 26SourcePDFScholar
2022

DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following

RA-L 2022

Language-guided Embodied AI benchmarks requiring an agent to navigate an environment and manipulate objects typically allow one-way communication: the human user gives a natural language command to the agent, and the agent can only follow the command passively. We present <bold xmlns:mml="http://www

Cited by 90SourcecodeScholar
2022

Learning to Act with Affordance-Aware Multimodal Neural SLAM

IROS 2022poster

Recent years have witnessed an emerging paradigm shift toward embodied artificial intelligence, in which an agent must learn to solve challenging tasks by interacting with its environment. There are several challenges in solving embodied multimodal tasks, including long-horizon planning, vision-and-…

Cited by 18SourcecodeScholar
2022

Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering

CVPR 2022poster

Outside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information to answer the question. Most previous works address the problem by first fusing the image and question in the multi-moda…

Cited by 113PDFcodeScholar