← Search

MengMeng Yang

16 accepted papers

2026

FASIONAD: Adaptive Uncertainty-Gated Fast–Slow Fusion Framework for Safe Autonomous Driving

ICRA 2026poster

Previous fast–slow system architectures demonstrated that pairing a reactive E2E planner with a deliberative vision-language model (VLM) can address these long-tail scenarios. However, these dual-system models that query the slow module at fixed intervals are computationally inefficient and introduc…

Cited by 0Scholar
2026

MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases

ICRA 2026poster

Vision-Language Models (VLMs) have demonstrated significant potential for end-to-end autonomous driving, yet a substantial gap remains between their current capabilities and the reliability necessary for real-world deployment. A critical challenge is their fragility, characterized by hallucinations …

2025

AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

EMNLP 2025

Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. To overcome this, we introduce AgentThink , a pioneering unified framewor

2025

C2F-Planner: Interaction-Aware Coarse-to-Fine Planning for Autonomous Vehicles

RA-L 2025

Ensuring safe and socially compliant driving is essential for autonomous vehicle planning. However, one of the significant challenges remains the performance bottleneck caused by interaction uncertainty in complex traffic scenarios. Traditional planning algorithms typically account for all traffic p

Cited by 0SourcecodeScholar
2025

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

NeurIPS 2025poster

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to sepa…

Cited by 0SourcecodeScholar
2025

EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Driving

IROS 2025

3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or tra

Cited by 7SourcecodeScholar
2025

Efficient End-to-end Visual Localization for Autonomous Driving with Decoupled BEV Neural Matching

IROS 2025

Accurate localization plays an important role in high-level autonomous driving systems. Conventional map matching-based localization methods solve the poses by explicitly matching map elements with sensor observations, generally sensitive to perception noise, therefore requiring costly hyperparamete

Cited by 1SourceScholar
2025

Enhancing Lane Segment Perception and Topology Reasoning With Crowdsourcing Trajectory Priors

RA-L 2025

In autonomous driving, recent advances in online mapping provide autonomous vehicles with a comprehensive understanding of driving scenarios. Moreover, incorporating prior information input into such perception model represents an effective approach to ensure the robustness and accuracy. However, ut

Cited by 3SourcecodeScholar
2025

LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction

IROS 2025

Accurate spatial and motion understanding is critical for autonomous driving systems. While object-level perception models excel in structured environments, they struggle with open-set categories and often lack precise geometric representation. Occupancy-based, class-agnostic methods offer better sc

Cited by 6SourceScholar
2025

Residual Learning Towards High-Fidelity Vehicle Dynamics Modeling With Transformer

RA-L 2025

The vehicle dynamics model serves as a vital component of autonomous driving systems, as it describes the temporal changes in vehicle state. Traditional physics-based methods employ mathematical formulae to model vehicle dynamics, but they are unable to adequately describe complex vehicle systems du

Cited by 6SourceScholar
2024

DiffMap: Enhancing Map Segmentation With Map Prior Using Diffusion Model

RA-L 2024

Constructing high-definition (HD) maps is a crucial requirement for enabling autonomous driving. In recent years, several map segmentation algorithms have been developed to address this need, leveraging advancements in Bird's-Eye View (BEV) perception. However, existing models still encounter challe

Cited by 17SourceScholar
2024

Poses as Queries: End-to-End Image-to-LiDAR Map Localization With Transformers

RA-L 2024

High-precision vehicle localization with commercial setups is a crucial technique for high-level autonomous driving tasks. As a newly emerged approach, monocular localization in LiDAR map achieves promising balance between cost and accuracy, but estimating pose by finding correspondences between suc

Cited by 8SourceScholar
2024

StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation

CVPR 2024highlight

Predicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict uniform snapshots of future occupancy states and require strictly…

2023

SGFNet: Segmentation Guided Fusion Network for 3D Object Detection

RA-L 2023

The self-driving application requires accurate 3D object detection as it is essential in several tasks, such as path and motion planning. However, up until this point, fusion-based detectors with cameras and LiDAR sensors have always been inferior to LiDAR-only detectors. This can be attributed to t

Cited by 4SourceScholar
2022

Attribute-Based Progressive Fusion Network for RGBT Tracking

AAAI 2022technical

RGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough an…

Cited by 171SourcePDFScholar
2022

Roadside HD Map Object Reconstruction Using Monocular Camera

RA-L 2022

The traditional HD map production based on mapping vehicles equipped with expensive sensors (e.g. LiDAR) struggles to keep the frequently changing HD maps up-to-date, thus calling for great research efforts in the HD map reconstruction using low-cost vision sensors. However, most existing works in t

Cited by 14SourceScholar