← Search

Rohan Choudhury

6 accepted papers

2026

FPS-Bench: A Benchmark for High Frame-Rate Video Understanding

CVPR 2026

Modern video-language models are typically trained on videos downsampled to low frames-per-second (FPS), and the most commonly used evaluation benchmarks are designed for low-FPS input as well. To address this shortcoming, we present FPS-Bench, a large video question-answering benchmark designed to

Cited by 0SourceScholar
2026

Faster Vision Transformers with Adaptive Patches

ICLR 2026poster

Vision Transformers (ViTs) partition input images into uniformly sized patches regardless of their content, resulting in long input sequence lengths for high-resolution images. We present Adaptive Patch Transformers (APT), which addresses this by using multiple different patch sizes within the same…

Cited by 0SourcecodeScholar
2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

AAAI 2026technical

We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial

Cited by 0SourcePDFScholar
2024

Don't Look Twice: Faster Video Transformers with Run-Length Tokenization

NeurIPS 2024spotlight

Video transformers are slow to train due to extremely large numbers of input tokens, even though many video tokens are repeated over time. Existing methods to remove uninformative tokens either have significant overhead, negating any speedup, or require tuning for different datasets and examples. We…

2024

JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle Interactions

ICRA 2024poster

Developing autonomous vehicles that can safely interact with pedestrians requires large amounts of pedestrian and vehicle data in order to learn accurate pedestrian-vehicle interaction models. However, gathering data that include crucial but rare scenarios - such as pedestrians jaywalking into heavy…

Cited by 0SourceScholar
2023

TEMPO: Efficient Multi-View Pose Estimation, Tracking, and Forecasting

ICCV 2023poster

Existing volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a robust spatiotemporal representation, improving pose accuracy w…

Cited by 26PDFcodeScholar