← Search

Aditya Chinchure

4 accepted papers

2026

SPIKE-RL: Video-LLMs meet Bayesian Surprise

ICLR 2026poster

Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing critical moments that define a video's narrative. We introduce SPIKE, an inference-time framework that quantifies Bayesian…

Cited by 0SourcecodeScholar
2025

Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events

CVPR 2025poster

The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios, making it difficult to discern whether model performance stems from keen perception and re…

Cited by 0SourcePDFScholar
2025

Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models

EMNLP 2025

The biases exhibited by text-to-image (TTI) models are often treated as independent, though in reality, they may be deeply interrelated. Addressing bias along one dimension—such as ethnicity or age—can inadvertently affect another, like gender, either mitigating or exacerbating existing disparities.

Cited by 0SourcePDFScholar
2024

From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models

EMNLP 2024main

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models’ cultural inclusivity. Still, they have limited coverage of cultures a…

Cited by 10SourcePDFScholar