← Search

Jingru Luo

2 accepted papers

2026

MTA: Multimodal Task Alignment for BEV Perception and Captioning

CVPR 2026

Bird's eye view (BEV)-based 3D perception plays a crucial role in autonomous driving applications. The rise of large language models has spurred interest in BEV-based captioning to understand object behavior in the surrounding environment. However, existing approaches treat perception and captioning

Cited by 6SourceScholar