← Search

Apratim Bhattacharyya

15 accepted papers

2026

Can Vision-Language Models Answer Face to Face Questions in the Real-World?

ICLR 2026poster

AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to converse with users in real-time using audio input. This raises the question: have we reached the point where AI models, c…

Cited by 0SourceScholar
2026

Enhancing Hallucination Detection through Noise Injection

ICLR 2026poster

Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as hallucinations. Effectively detecting hallucinations is therefore crucial for the safe deployment of LLMs. Recent research has linked hallucinations to model uncertainty, suggesting that hallucinations…

Cited by 0SourceScholar
2026

Notes-To-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks

ICRA 2026poster

Many dexterous manipulation tasks are non-markovian in nature, yet little attention has been paid to this fact in the recent upsurge of the vision-language-action (VLA) paradigm. Although they are successful in bringing internet-scale semantic understanding to robotics, existing VLAs are primarily "…

2026

RoCA: Robust Cross-Domain End-to-End Autonomous Driving

ICML 2026poster

End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deployment across domains (e.g., cities). Although several works have incorporated Large Language Models (LLMs) to leverage the…

Cited by 0SourceScholar
2025

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

NeurIPS 2025poster

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, a…

Cited by 0SourcecodeScholar
2025

Distilling Multi-modal Large Language Models for Autonomous Driving

CVPR 2025poster

Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational cos…

Cited by 4SourcePDFScholar
2024

ClevrSkills: Compositional Language And Visual Reasoning in Robotics

NeurIPS 2024poster

Robotics tasks are highly compositional by nature. For example, to perform a high-level task like cleaning the table a robot must employ low-level capabilities of moving the effectors to the objects on the table, pick them up and then move them off the table one-by-one, while re-evaluating the conse…

2024

Look, Remember and Reason: Grounded Reasoning in Videos with Language Models

ICLR 2024poster

Multi-modal language models (LM) have recently shown promising performance in high-level reasoning tasks on videos. However, existing methods still fall short in tasks like causal or compositional spatiotemporal reasoning over actions, in which model predictions need to be grounded in fine-grained…

Cited by 18SourcePDFScholar
2024

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction

NeurIPS 2024poster

Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e., prompted) by the user. Open-ended, asynchronous interactions, where an AI model may proactively deliver timely respon…

2022

KING: Generating Safety-Critical Driving Scenarios for Robust Imitation via Kinematics Gradients

ECCV 2022poster

"Simulators offer the possibility of safe, low-cost development of self-driving systems. However, current driving simulators exhibit naïve behavior models for background traffic. Hand-tuned scenarios are typically added during simulation to induce safety-critical situations. An alternative approach…

2021

Euro-PVI: Pedestrian Vehicle Interactions in Dense Urban Centers

CVPR 2021poster

Accurate prediction of pedestrian and bicyclist paths is integral to the development of reliable autonomous vehicles in dense urban environments. The interactions between vehicle and pedestrian or bicyclist have a significant impact on the trajectories of traffic participants e.g. stopping or turnin…

Cited by 45PDFScholar
2020

Normalizing Flows With Multi-Scale Autoregressive Priors

CVPR 2020poster

Flow-based generative models are an important class of exact inference models that admit efficient inference and sampling for image synthesis. Owing to the efficiency constraints on the design of the flow layers, e.g. split coupling flow layers in which approximately half the pixels do not undergo f…

Cited by 14PDFcodeScholar
2019

Bayesian Prediction of Future Street Scenes using Synthetic Likelihoods

ICLR 2019poster

For autonomous agents to successfully operate in the real world, the ability to anticipate future scene states is a key competence. In real-world scenarios, future states become increasingly uncertain and multi-modal, particularly on long time horizons. Dropout based Bayesian inference provides a co…

Cited by 56SourcePDFScholar
2018

Accurate and Diverse Sampling of Sequences Based on a “Best of Many” Sample Objective

CVPR 2018poster

For autonomous agents to successfully operate in the real world, anticipation of future events and states of their environment is a key competence. This problem has been formalized as a sequence extrapolation problem, where a number of observations are used to predict the sequence into the future. R…

Cited by 140SourcePDFScholar
2018

Long-Term On-Board Prediction of People in Traffic Scenes Under Uncertainty

CVPR 2018poster

Progress towards advanced systems for assisted and autonomous driving is leveraging recent advances in recognition and segmentation methods. Yet, we are still facing challenges in bringing reliable driving to inner cities, as those are composed of highly dynamic scenes observed from a moving platfo…

Cited by 288SourcePDFScholar