← Search

Liuyi Wang

13 accepted papers

2026

BEVDrive-E2E: Imitation With Bird's Eye View Perception for Interpretable End-to-End Autonomous Driving

RA-L 2026

Imitation learning (IL) for end-to-end autonomous driving (E2E-AD) has made great progress recently in the closed-loop evaluation of the CARLA simulator. However, the causal confusion remains an open problem. To address this issue, we propose the BEVDrive-E2E to explore the interpretability of the e

Cited by 0SourceScholar
2026

GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding

CVPR 2026

Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame sampling to extract video information, resulting in a sparse distribution of ke

Cited by 0SourcecodeScholar
2026

TerrFlat: Physics-Driven Geometry Representation for Structure-Aware Freespace Detection

ICRA 2026poster

Freespace detection in autonomous driving is limited by the lack of explicit geometric modeling, hindering generalization across complex terrains. Existing approaches are predominantly data-driven and neglect the physical structure of drivable surfaces. We propose Terrain Flat (TerrFlat), a physics-…

Cited by 0SourceScholar
2025

CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation

ICCV 2025poster

In the effort to achieve robust and generalizable category-level object pose estimation, recent methods primarily focus on learning fundamental representations from data. However, the inherent biases within the data are often overlooked: the repeated training samples and similar environments may mis…

2025

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

CVPR 2025poster

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and s…

2025

Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities

ICCV 2025poster

Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we introduce VLN-PE, a physically realistic VLN platform supporting humanoid, quadru…

2024

Enhanced Language-guided Robot Navigation with Panoramic Semantic Depth Perception and Cross-modal Fusion

IROS 2024poster

Integrating visual observation with linguistic instruction holds significant promise for enhancing robot navigation across unstructured environments and enriches the human-robot interaction experience. However, while panoramic RGB views furnish robots with extensive environmental visuals, current me…

Cited by 0SourcecodeScholar
2024

Multimodal Evolutionary Encoder for Continuous Vision-Language Navigation

IROS 2024poster

Can multimodal encoder evolve when facing increasingly tough circumstances? Our work investigates this possibility in the context of continuous vision-language navigation (continuous VLN), which aims to navigate robots under linguistic supervision and visual feedback. We propose a multimodal evoluti…

Cited by 0SourcecodeScholar
2024

Vision-and-Language Navigation via Causal Learning

CVPR 2024poster

In the pursuit of robust and generalizable environment perception and language understanding the ubiquitous challenge of dataset bias continues to plague vision-and-language navigation (VLN) agents hindering their performance in unseen environments. This paper introduces the generalized cross-modal…

2023

A Dual Semantic-Aware Recurrent Global-Adaptive Network for Vision-and-Language Navigation

IJCAI 2023poster

Vision-and-Language Navigation (VLN) is a realistic but challenging task that requires an agent to locate the target region using verbal and visual cues. While significant advancements have been achieved recently, there are still two broad limitations: (1) The explicit information mining for signifi…

2023

Multiple Thinking Achieving Meta-Ability Decoupling for Object Navigation

ICML 2023poster

We propose a meta-ability decoupling (MAD) paradigm, which brings together various object navigation methods in an architecture system, allowing them to mutually enhance each other and evolve together. Based on the MAD paradigm, we design a multiple thinking (MT) model that leverages distinct thinki…

Cited by 10SourcePDFScholar
2023

Search for or Navigate to? Dual Adaptive Thinking for Object Navigation

ICCV 2023poster

"Search for" or "Navigate to"? When we find a specific object in an unknown environment, the two choices always arise in our subconscious mind. Before we see the target, we search for the target based on prior experience. Once we have seen the target, we can navigate to it by remembering the target…

Cited by 21PDFScholar