← Search

Yanyuan Qiao

16 accepted papers

2026

Fast-SmartWay: Panoramic-Free End-To-End Zero-Shot Vision-And-Language Navigation

ICRA 2026poster

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panoramic observations and two-stage pipelines involving waypoint predictors, which i…

2025

A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making

EMNLP 2025

Medical decision-making often involves integrating knowledge from multiple clinical specialties, typically achieved through multidisciplinary teams. Inspired by this collaborative process, recent work has leveraged large language models (LLMs) in multi-agent collaboration frameworks to emulate exper

2025

COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation

ICCV 2025poster

Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional component…

2025

Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation

NeurIPS 2025poster

We present Chain-of-Action (CoA), a novel visuomotor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-le…

Cited by 0SourceScholar
2025

Effective Tuning Strategies for Generalist Robot Manipulation Policies

ICRA 2025

Generalist robot manipulation policies (GMPs) have the potential to generalize across a wide range of tasks, devices, and environments. However, existing policies continue to struggle with out-of-distribution scenarios due to the inherent difficulty of collecting sufficient action data to cover exte

Cited by 9SourceScholar
2025

General Scene Adaptation for Vision-and-Language Navigation

ICLR 2025poster

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner. However, real-world navigation robots often operate in pers…

2025

Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments

ICRA 2025

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, dealing with visually diverse scenes or transitioning from simulated environments to real-world deployment is still challenging.

Cited by 7SourceScholar
2025

MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation

ICRA 2025

In recent years, Embodied Artificial Intelligence (Embodied AI) has advanced rapidly, yet the increasing size of models conflicts with the limited computational capabilities of Embodied AI platforms. To address this challenge, we aim to achieve both high model performance and practical deployability

Cited by 5SourceScholar
2025

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero…

Cited by 0SourceScholar
2025

Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs

ICRA 2025

Vision-and-Language Navigation (VLN) tasks require an agent to follow textual instructions to navigate through 3D environments. Traditional approaches use supervised learning methods, relying heavily on domain-specific datasets to train VLN models. Recent methods try to utilize closedsource large la

Cited by 49SourceScholar
2025

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

IROS 2025

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint predictor to generate waypoints and a navigator to execute movement

Cited by 13SourceScholar
2025

Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs

EMNLP 2025

Integrating large language models (LLMs) into embodied AI models is becoming increasingly prevalent. However, existing zero-shot LLM-based Vision-and-Language Navigation (VLN) agents either encode images as textual scene descriptions, potentially oversimplifying visual details, or process raw image

2023

March in Chat: Interactive Prompting for Remote Embodied Referring Expression

ICCV 2023poster

Many Vision-and-Language Navigation (VLN) tasks have been proposed in recent years, from room-based to object-based and indoor to outdoor. The REVERIE (Remote Embodied Referring Expression) is interesting since it only provides high-level instructions to the agent, which are closer to human commands…

Cited by 39PDFcodeScholar
2022

HOP: History-and-Order Aware Pre-Training for Vision-and-Language Navigation

CVPR 2022poster

Pre-training has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, previous pre-training methods for VLN either lack the ability to predict future actions or ignore the trajectory contexts, which are essential for a greedy navigation process. In this work, to p…

Cited by 94PDFcodeScholar