← Search

Meng Tian

7 accepted papers

2026

Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning

AAAI 2026technical

Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by relying heavily on history input information, achieving seemi

Cited by 0SourcePDFScholar
2026

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

CVPR 2026

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak at spatial grounding and understanding, and VLA systems bui

Cited by 0SourceScholar
2025

Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

ICCV 2025poster

Existing benchmarks for Vision-Language Model (VLM) in autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce VLAD…

2025

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPS 2025spotlight

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual features and insufficient real-time spatiotemporal reasoning. To a…

Cited by 0SourceScholar
2024

Exploring Spatio-Temporal Discriminative Cues for Group Activity Recognition Via Contrastive Learning

ICASSP 2024accepted

Group activity recognition is a challenging task that involves multiple moving actors within a cluttered scene. Existing methods often rely on object detector to avoid individual bounding box labeling during testing, but are prone to false detections due to factors such as occlusion and background c…

Cited by 0SourceScholar
2020

Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation

ECCV 2020poster

We present a novel learning approach to recover the 6D poses and sizes of unseen object instances from an RGB-D image. To handle the intra-class shape variation, we propose a deep network to reconstruct the 3D object model by explicitly modeling the deformation from a pre-learned categorical shape p…