← Search

Yibo Cui

2 accepted papers

2023

Grounded Entity-Landmark Adaptive Pre-Training for Vision-and-Language Navigation

ICCV 2023oral

Cross-modal alignment is one key challenge for Vision-and-Language Navigation (VLN). Most existing studies concentrate on mapping the global instruction or single sub-instruction to the corresponding trajectory. However, another critical problem of achieving fine-grained alignment at the entity leve…

Cited by 21PDFcodeScholar