← Search

Oleg Sinavski

3 accepted papers

2025

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

CVPR 2025highlight

Integrating large language models (LLMs) into autonomous driving has attracted significant attention with the hope of improving generalization and explainability. However, existing methods often focus on either driving or vision-language understanding but achieving both high driving performance and…

Cited by 1SourcePDFScholar
2024

Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

ICRA 2024poster

Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique objectlevel multimodal LLM architecture that merges vectorized numeric modalities with a pre-trained LLM to improve context understanding in dr…

Cited by 229SourcecodeScholar
2024

LingoQA: Video Question Answering for Autonomous Driving

ECCV 2024poster

"We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capab…