← Search

Ahmadreza Moradipari

6 accepted papers

2025

NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models

ICCV 2025poster

Recent advances in multi-modal large language models (MLLMs) have demonstrated strong performance across various domains; however, their ability to comprehend driving scenes remains less proven. The complexity of driving scenarios, which includes multi-view information, poses significant challenges…

2025

On Learning Closed-Loop Probabilistic Multi-Agent Simulator

IROS 2025

The rapid iteration of autonomous vehicle (AV) deployments leads to increasing needs for building realistic and scalable multi-agent traffic simulators for efficient evaluation. Recent advances in this area focus on closed-loop simulators that enable generating diverse and interactive scenarios. Thi

Cited by 0SourceScholar
2023

Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning

NeurIPS 2023poster

In this paper, we prove state-of-the-art Bayesian regret bounds for Thompson Sampling in reinforcement learning in a multitude of settings. We present a refined analysis of the information ratio, and show an upper bound of order $\widetilde{O}(H\sqrt{d_{l_1}T})$ in the time inhomogeneous reinforceme…

Cited by 5SourcePDFScholar
2022

Feature and Parameter Selection in Stochastic Linear Bandits

ICML 2022spotlight

We study two model selection settings in stochastic linear bandits (LB). In the first setting, which we refer to as feature selection, the expected reward of the LB problem is in the linear span of at least one of $M$ feature maps (models). In the second setting, the reward parameter of the LB probl…

Cited by 11SourcePDFScholar
2020

Linear Thompson Sampling Under Unknown Linear Constraints

ICASSP 2020accepted

We study how adding unknown linear safety constraints affects the performance of Thompson Sampling in the linear stochastic bandit problem. The additional constraints must be met at each round in spite of uncertainty about the environment requiring that the learner acts conservatively in choosing he…

Cited by 0SourceScholar