← Search

Ramin Mehran

3 accepted papers

2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

CVPR 2026

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of inter- mediate reasoning steps, and most provide answers only in the text doma

Cited by 0SourcecodeScholar
2025

MINERVA: Evaluating Complex Video Reasoning

ICCV 2025poster

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able to combine perceptual and temporal information to reason ab…

2020

Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing

ICASSP 2020accepted

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a…

Cited by 0SourceScholar