← Search

Shyamal Buch

15 accepted papers

2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

CVPR 2026

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of inter- mediate reasoning steps, and most provide answers only in the text doma

Cited by 0SourcecodeScholar
2025

Flexible Frame Selection for Efficient Video Reasoning

CVPR 2025poster

Video-language models have shown promise for addressing a range of multimodal tasks for video understanding, such as video question-answering. However, the inherent computational challenges of processing long video data and increasing model sizes have led to standard approaches that are limited by t…

Cited by 0SourcePDFScholar
2025

MINERVA: Evaluating Complex Video Reasoning

ICCV 2025poster

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able to combine perceptual and temporal information to reason ab…

2025

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

EMNLP 2025

Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for evaluating longer, more diverse captions and in settings where ground-truth human-annotated captions are unavailable. We in

Cited by 0SourcePDFScholar
2024

Mixture of Nested Experts: Adaptive Processing of Visual Tokens

NeurIPS 2024poster

The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. While Vision Transformer (ViT) based models scale effectively to large data regimes, they fail to capitalize on this inher…

Cited by 8SourcePDFScholar
2024

MoReVQA: Exploring Modular Reasoning Models for Video Question Answering

CVPR 2024poster

This paper addresses the task of video question answering (videoQA) via a decomposed multi-stage modular reasoning framework. Previous modular methods have shown promise with a single planning stage ungrounded in visual content. However through a simple and effective baseline we find that such syste…

Cited by 32SourcePDFScholar
2024

Streaming Dense Video Captioning

CVPR 2024poster

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos predict rich detailed textual descriptions and be able to produce outputs before processing the entire video. Current state-of-the-art models however process…

2024

Streaming Detection of Queried Event Start

NeurIPS 2024poster

Robotics, autonomous driving, augmented reality, and many embodied computer vision applications must quickly react to user-defined events unfolding in real time. We address this setting by proposing a novel task for multimodal video understanding---Streaming Detection of Queried Event Start (SDQES).…

2022

Revisiting the "Video" in Video-Language Understanding

CVPR 2022oral

What makes a video task uniquely suited for videos, beyond what can be understood from a single image? Building on recent progress in self-supervised image-language models, we revisit this question in the context of video and language tasks. We propose the atemporal probe (ATP), a new model for vide…

Cited by 199PDFScholar
2021

BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments

CoRL 2021poster

We introduce BEHAVIOR, a benchmark for embodied AI with 100 activities in simulation, spanning a range of everyday household chores such as cleaning, maintenance, and food preparation. These activities are designed to be realistic, diverse and complex, aiming to reproduce the challenges that agents…

Cited by 176SourceScholar
2021

iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes

IROS 2021poster

We present iGibson 1.0, a novel simulation environment to develop robotic solutions for interactive tasks in large-scale realistic scenes. Our environment contains 15 fully interactive home-sized scenes with 108 rooms populated with rigid and articulated objects. The scenes are replicas of real-worl…

Cited by 193SourceScholar
2020

RubiksNet: Learnable 3D-Shift for Efficient Video Action Recognition

ECCV 2020poster

Video action recognition is a complex task dependent on modeling spatial and temporal context. Standard approaches rely on 2D or 3D convolutions to process such context, resulting in expensive operations with millions of parameters. Recent efficient architectures leverage a channel-wise shift-based…

2018

End-to-End Joint Semantic Segmentation of Actors and Actions in Video

ECCV 2018poster

Traditional video understanding tasks include human action recognition and actor/object semantic segmentation. However, the combined task of providing semantic segmentation for different actor classes simultaneously with their action class remains a challenging but necessary task for many applicatio…

Cited by 52SourcePDFScholar
2018

Finding "It": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos

CVPR 2018poster

Grounding textual phrases in visual content with standalone image-sentence pairs is a challenging task. When we consider grounding in instructional videos, this problem becomes profoundly more complex: the latent temporal structure of instructional videos breaks independence assumptions and necessit…

Cited by 114SourcePDFScholar
2017

SST: Single-Stream Temporal Action Proposals

CVPR 2017poster

Our paper presents a new approach for temporal detection of human actions in long, untrimmed video sequences. We introduce Single-Stream Temporal Action Proposals (SST), a new effective and efficient deep architecture for the generation of temporal action proposals. Our network can run continuously…

Cited by 544PDFcodeScholar