← Search

Wenzhe Cai

16 accepted papers

2026

Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation

ICLR 2026poster

While recent large vision-language models (VLMs) have improved generalization in vision-language navigation (VLN), existing methods typically rely on end-to-end pipelines that map vision-language inputs directly to short-horizon discrete actions. Such designs often produce fragmented motions, incur…

Cited by 0SourceScholar
2026

LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory

ICML 2026poster

Automated laboratories hold the promise of accelerating scientific discovery, yet their deployment is bottlenecked by the difficulty of designing safe and executable environments. While simulator-based design offers scalability, existing 3D scene generation methods are primarily tailored for househo…

Cited by 0SourceScholar
2026

LoGoPlanner: Localization Grounded Navigation Policy with Metric-Aware Visual Geometry

ICRA 2026poster

Trajectory planning in unstructured environments is a fundamental and challenging capability for mobile robots. Traditional modular pipelines suffer from latency and cascading errors across perception, localization, mapping, and planning modules. Recent end-to-end learning methods map raw visual obs…

2026

NavDP: Learning Sim-To-Real Navigation Diffusion Policy with Privileged Information Guidance

ICRA 2026poster

Learning to navigate in dynamic and complex open-world environments is a critical yet challenging capability for autonomous robots. Existing approaches often rely on cascaded modular frameworks, which require extensive hyperparameter tuning or learning from limited real-world demonstration data. In …

2026

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

ICRA 2026poster

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reasoning capabilities. In this work, we introduce the NavSpace benchmark, which conta…

2026

Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP

ICRA 2026poster

Object-oriented embodied navigation tasks require agents to locate specific objects, either defined by category or images, in unseen environments. While recent methods have made progress in extending closed-set models to open-vocabulary scenarios with foundation models, they typically rely on traini…

Cited by 0Scholar
2026

StreamVLN: Streaming Vision-And-Language Navigation Via SlowFast Context Modeling

ICRA 2026poster

Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs) have driven recent progress, current VLN methods based on Vid…

2025

Boosting Efficient Reinforcement Learning for Vision-and-Language Navigation With Open-Sourced LLM

RA-L 2025

Vision-and-Language Navigation (VLN) requires an agent to navigate in photo-realistic environments based on language instructions. Existing methods typically employ imitation learning to train agents. However, approaches based on recurrent neural networks suffer from poor generalization, while trans

Cited by 13SourceScholar
2025

ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

ICLR 2025poster

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLM…

Cited by 2SourcePDFScholar
2025

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

NeurIPS 2025poster

The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collis…

Cited by 0SourceScholar
2024

Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill

ICRA 2024poster

Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the locomotion part, most works still depend on map-based planning appro…

Cited by 39SourcecodeScholar
2024

Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

ICRA 2024poster

Visual language navigation (VLN) is an embodied task demanding a wide range of skills encompassing understanding, perception, and planning. For such a multifaceted challenge, previous VLN methods totally rely on one model’s own thinking to make predictions within one round. However, existing models,…

Cited by 51SourceScholar
2024

InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

CoRL 2024poster

Enabling robots to navigate following diverse language instructions in unexplored environments is an attractive goal for human-robot interaction. However, this goal is challenging because different navigation tasks require different strategies. The scarcity of instruction navigation data hinders tra…

Cited by 34SourceScholar
2024

MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven Navigation

NeurIPS 2024poster

The process of satisfying daily demands is a fundamental aspect of humans' daily lives. With the advancement of embodied AI, robots are increasingly capable of satisfying human demands. Demand-driven navigation (DDN) is a task in which an agent must locate an object to satisfy a specified demand ins…

Cited by 0SourcePDFScholar
2023

Robust Navigation with Cross-Modal Fusion and Knowledge Transfer

ICRA 2023poster

Recently, learning-based approaches show promising results in navigation tasks. However, the poor generalization capability and the simulation-reality gap prevent a wide range of applications. We consider the problem of improving the generalization of mobile robots and achieving sim-to-real transfer…

Cited by 1SourcecodeScholar