← Search

Qiaozi Gao

8 accepted papers

2024

GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

CVPR 2024poster

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level representations that are important for fine-grained visual understa…

Cited by 47SourcePDFScholar
2024

Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuning

ICML 2024poster

Prompt-based learning has been demonstrated as a compelling paradigm contributing to large language models' tremendous success (LLMs). Inspired by their success in language tasks, existing research has leveraged LLMs in embodied instruction following and task planning. In this work, we tackle the pr…

Cited by 11SourcePDFScholar
2023

Alexa Arena: A User-Centric Interactive Platform for Embodied AI

NeurIPS 2023poster

We introduce Alexa Arena, a user-centric simulation platform to facilitate research in building assistive conversational embodied agents. Alexa Arena features multi-room layouts and an abundance of interactable objects. With user-friendly graphics and control mechanisms, the platform supports the de…

2023

LEMMA: Learning Language-Conditioned Multi-Robot Manipulation

RA-L 2023

Complex manipulation tasks often require robots with complementary capabilities to collaborate. We introduce a benchmark for <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">L</u> anguag <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xml

Cited by 15SourceScholar
2022

DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following

RA-L 2022

Language-guided Embodied AI benchmarks requiring an agent to navigate an environment and manipulate objects typically allow one-way communication: the human user gives a natural language command to the agent, and the agent can only follow the command passively. We present <bold xmlns:mml="http://www

Cited by 90SourcecodeScholar
2022

Learning to Act with Affordance-Aware Multimodal Neural SLAM

IROS 2022poster

Recent years have witnessed an emerging paradigm shift toward embodied artificial intelligence, in which an agent must learn to solve challenging tasks by interacting with its environment. There are several challenges in solving embodied multimodal tasks, including long-horizon planning, vision-and-…

Cited by 18SourcecodeScholar
2022

Towards Large-Scale Interpretable Knowledge Graph Reasoning for Dialogue Systems

ACL 2022findings

Users interacting with voice assistants today need to phrase their requests in a very specific manner to elicit an appropriate response. This limits the user experience, and is partly due to the lack of reasoning capabilities of dialogue platforms and the hand-crafted rules that require extensive la…

2021

Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding

EMNLP 2021finding

Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks. However, evaluations only based on end task performance shed little light on machines’ true ability in language understanding and reasoning. In this paper, we highlight…