← Search

Dibyadip Chatterjee

6 accepted papers

2026

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

ICML 2026poster

Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and evicting old key-value(KV) caches is required for unbounded streams. Secondly, due to the high cost of collecting and tra…

Cited by 0SourceScholar
2026

On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have advanced open-world action understanding and can be adapted as generative classifiers for closed-set settings by autoregressively generating action labels as text. However, this approach is inefficient, and shared subwords across action labels introduce…

Cited by 0SourcecodeScholar
2025

Streaming VideoLLMs for Real-Time Procedural Video Understanding

ICCV 2025poster

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens -- verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded…

Cited by 9SourcePDFScholar
2024

On the Utility of 3D Hand Poses for Action Recognition

ECCV 2024poster

"3D hand pose is an underexplored modality for action recognition. Poses are compact yet informative and can greatly benefit applications with limited compute budgets. However, poses alone offer an incomplete understanding of actions, as they cannot fully capture objects and environments with which…

2023

Opening the Vocabulary of Egocentric Actions

NeurIPS 2023poster

Human actions in egocentric videos often feature hand-object interactions composed of a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations — sparsity of action compositions and a closed set of interacting objects. Thi…

2022

Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

CVPR 2022poster

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101…

Cited by 246PDFcodeScholar