← Search

Xiwen Liang

11 accepted papers

2025

PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly

NeurIPS 2025poster

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressiv…

Cited by 0SourceScholar
2025

Structured Preference Optimization for Vision-Language Long-Horizon Task Planning

EMNLP 2025

Existing vision-language planning methods perform well on short-horizon tasks but struggle with long-horizon reasoning in dynamic environments due to the difficulty of training models to generate high-quality reasoning processes. To address this, we propose Structured Preference Optimization (SPO),

Cited by 0SourcePDFScholar
2024

CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation

ACL 2024findings

Understanding and following natural language instructions while navigating through complex, real-world environments poses a significant challenge for general-purpose robots. These environments often include obstacles and pedestrians, making it essential for autonomous agents to possess the capabilit…

2023

NLIP: Noise-Robust Language-Image Pre-training

AAAI 2023technical

Large-scale cross-modal pre-training paradigms have recently shown ubiquitous success on a wide range of downstream tasks, e.g., zero-shot classification, retrieval and image captioning. However, their successes highly rely on the scale and quality of web-crawled data that naturally contain much inc…

Cited by 33SourcePDFScholar
2023

Visual Exemplar Driven Task-Prompting for Unified Perception in Autonomous Driving

CVPR 2023poster

Multi-task learning has emerged as a powerful paradigm to solve a range of tasks simultaneously with good efficiency in both computation resources and inference time. However, these algorithms are designed for different tasks mostly not within the scope of autonomous driving, thus making it hard to…

Cited by 21SourcePDFScholar
2022

ADAPT: Vision-Language Navigation With Modality-Aligned Action Prompts

CVPR 2022poster

Vision-Language Navigation (VLN) is a challenging task that requires an embodied agent to perform action-level modality alignment, i.e., make instruction-asked actions sequentially in complex visual environments. Most existing VLN agents learn the instruction-path data directly and cannot sufficient…

Cited by 58PDFScholar
2022

Contrastive Instruction-Trajectory Learning for Vision-Language Navigation

AAAI 2022technical

The vision-language navigation (VLN) task requires an agent to reach a target with the guidance of natural language instruction. Previous works learn to navigate step-by-step following an instruction. However, these works may fail to discriminate the similarities and discrepancies across instruction…

2022

Effective Adaptation in Multi-Task Co-Training for Unified Autonomous Driving

NeurIPS 2022accept

Aiming towards a holistic understanding of multiple downstream tasks simultaneously, there is a need for extracting features with better transferability. Though many latest self-supervised pre-training methods have achieved impressive performance on various vision tasks under the prevailing pretrain…

Cited by 38SourcePDFScholar
2022

Visual-Language Navigation Pretraining via Prompt-based Environmental Self-exploration

ACL 2022long

Vision-language navigation (VLN) is a challenging task due to its large searching space in the environment. To address this problem, previous works have proposed some methods of fine-tuning a large model that pretrained on large-scale datasets. However, the conventional fine-tuning methods require e…

2021

SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving

NeurIPS 2021poster

Aiming at facilitating a real-world, ever-evolving and scalable autonomous driving system, we present a large-scale dataset for standardizing the evaluation of different self-supervised and semi-supervised approaches by learning from raw data, which is the first and largest dataset to date. Existing…

Cited by 82SourcecodeScholar
2021

SOON: Scenario Oriented Object Navigation With Graph-Based Exploration

CVPR 2021poster

The ability to navigate like a human towards a language-guided target from anywhere in a 3D embodied environment is one of the 'holy grail' goals of intelligent robots. Most visual navigation benchmarks, however, focus on navigating toward a target from a fixed starting point, guided by an elaborate…

Cited by 131PDFcodeScholar