← Search

Mosam Dabhi

6 accepted papers

2026

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

AAAI 2026technical

We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial

Cited by 0SourcePDFScholar
2026

Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera

CVPR 2026

Modern perception increasingly relies on fisheye, panoramic, and other wide field-of-view (FoV) cameras, yet most pipelines still apply planar CNNs designed for pinhole imagery on 2D grids, where pixel-space neighborhoods misrepresent physical adjacency and models are sensitive to global rotations.

Cited by 0SourceScholar
2022

MBW: Multi-view Bootstrapping in the Wild

NeurIPS 2022accept

Labeling articulated objects in unconstrained settings has a wide variety of applications including entertainment, neuroscience, psychology, ethology, and many fields of medicine. Large offline labeled datasets do not exist for all but the most common articulated object categories (e.g., humans). Ha…

Cited by 2SourcePDFScholar
2019

Real-Time Information-Theoretic Exploration with Gaussian Mixture Model Maps

RSS 2019poster

This paper develops an exploration framework that leverages Gaussian mixture models (GMMs) for high-fidelity perceptual modeling and exploits the compactness of the distributions for information sharing in communications-constrained applications. State-of-the-art, high-resolution perceptual modeling…

Cited by 39SourcePDFScholar