← Search

Pierre-Louis Guhur

6 accepted papers

2022

Instruction-driven history-aware policies for robotic manipulations

CoRL 2022oral

In human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions. Yet, robotic manipulation is extremely challenging as it requires fine-grained motor control, long-term memory as well as generalization to previously unseen tasks and…

Cited by 117SourcecodeScholar
2022

Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

NeurIPS 2022accept

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as "the left most chair" and "a chair next to the window". In this work we propose a la…

2022

Learning from Unlabeled 3D Environments for Vision-and-Language Navigation

ECCV 2022poster

"In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training data, resulting in unsatisfactory generalization to unseen env…

2022

Think Global, Act Local: Dual-Scale Graph Transformer for Vision-and-Language Navigation

CVPR 2022oral

Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual scenes, but also should explore the environment to reach its target. In this work, we propose a dual-scale graph transfor…

Cited by 181PDFScholar
2021

Airbert: In-Domain Pretraining for Vision-and-Language Navigation

ICCV 2021poster

Vision-and-language navigation (VLN) aims to enable embodied agents to navigate in realistic environments using natural language instructions. Given the scarcity of domain-specific training data and the high diversity of image and language inputs, the generalization of VLN agents to unseen environme…

Cited by 169PDFcodeScholar
2021

History Aware Multimodal Transformer for Vision-and-Language Navigation

NeurIPS 2021poster

Vision-and-language navigation (VLN) aims to build autonomous visual agents that follow instructions and navigate in real scenes. To remember previously visited locations and actions taken, most approaches to VLN implement memory using recurrent states. Instead, we introduce a History Aware Multimod…

Cited by 268SourcePDFScholar