AAAI 2026technical0 citations

VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness

Qimao Chen, Fang Li, Shaoqing Xu, Zhiyi Lai, Zixun Xie, Yuechen Luo, Shengyin Jiang, Hanbing Li

Abstract

The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents

BibTeX
@inproceedings{aaai2026_viltaavlminthelo,
  title = {VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness},
  author = {Qimao Chen and Fang Li and Shaoqing Xu and Zhiyi Lai and Zixun Xie and Yuechen Luo and Shengyin Jiang and Hanbing Li and Long Chen and Bing Wang and Yi Zhang and Zhi-Xin Yang},
  booktitle = {AAAI 2026},
  year = {2026}
}
VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness · AAAI 2026