← Search

Sahithya Ravi

9 accepted papers

2026

SPIKE-RL: Video-LLMs meet Bayesian Surprise

ICLR 2026poster

Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing critical moments that define a video's narrative. We introduce SPIKE, an inference-time framework that quantifies Bayesian…

Cited by 0SourcecodeScholar
2025

Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events

CVPR 2025poster

The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios, making it difficult to discern whether model performance stems from keen perception and re…

Cited by 0SourcePDFScholar
2025

CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs’ Cultural Knowledge Through Human-AI Red-Teaming

ACL 2025long

Robust, diverse, and challenging cultural knowledge benchmarks are essential for measuring our progress towards making LMs that are helpful across diverse cultures. We introduce CulturalBench: a set of 1,696 human-written and human-verified questions to assess LMs’ cultural knowledge, covering 45 gl…

Cited by 0SourcePDFScholar
2025

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

ACL 2025finding

A person’s demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences…

Cited by 0SourcePDFScholar
2025

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

EMNLP 2025

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce Disjoint-3DQA , a generative QA benchmark that evaluates this ability

Cited by 0SourcePDFScholar
2024

From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models

EMNLP 2024main

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models’ cultural inclusivity. Still, they have limited coverage of cultures a…

Cited by 10SourcePDFScholar
2024

Small But Funny: A Feedback-Driven Approach to Humor Distillation

ACL 2024long

The emergence of Large Language Models (LLMs) has brought to light promising language generation capabilities, particularly in performing tasks like complex reasoning and creative writing. Consequently, distillation through imitation of teacher responses has emerged as a popular technique to transfe…

Cited by 4SourcePDFScholar
2023

CASE: Commonsense-Augmented Score with an Expanded Answer Space

EMNLP 2023long findings

LLMs have demonstrated impressive zero-shot performance on NLP tasks thanks to the knowledge they acquired in their training. In multiple-choice QA tasks, the LM probabilities are used as an imperfect measure of the plausibility of each answer choice. One of the major limitations of the basic score…

Cited by 0SourcecodeScholar
2023

COMET-M: Reasoning about Multiple Events in Complex Sentences

EMNLP 2023long findings

Understanding the speaker’s intended meaning often involves drawing commonsense inferences to reason about what is not stated explicitly. In multi-event sentences, it requires understanding the relationships between events based on contextual knowledge. We propose COMET-M (Multi-Event), an event-cen…

Cited by 0SourcecodeScholar