← Search

Chonghao Sima*

1 accepted papers

2024

DriveLM: Driving with Graph Visual Question Answering

ECCV 2024oral

"We study how vision-language models (VLMs) trained on web-scale data can be integrated into end-to-end driving systems to boost generalization and enable interactivity with human users. While recent approaches adapt VLMs to driving via single-round visual question answering (VQA), human drivers rea…