← Search

Ehsan Elhamifar

33 accepted papers

2026

AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models

CVPR 2026

Virtual task assistants must recognize and explain users' mistakes to provide effective and corrective guidance. In this paper, we address the problem of error reasoning in long task videos, which is to detect and explain errors. Although recent Vision-Language Models (VLMs) demonstrate strong capab

Cited by 0SourcecodeScholar
2026

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding

CVPR 2026

Large vision-language models struggle with medical video understanding, where spatial precision, temporal reasoning, and clinical semantics are critical. To address this, we first introduce MedVidBench, a large-scale benchmark of 531,850 video-instruction pairs across 8 medical sources spanning vide

Cited by 0SourceScholar
2025

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

CVPR 2025poster

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challen…

2024

Error Detection in Egocentric Procedural Task Videos

CVPR 2024poster

We present a new egocentric procedural error dataset containing videos with various types of errors as well as normal videos and propose a new framework for procedural error detection using error-free training videos only. Our framework consists of an action segmentation model and a contrastive step…

Cited by 15SourcePDFScholar
2024

FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action Segmentation

CVPR 2024poster

We study supervised action segmentation whose goal is to predict framewise action labels of a video. To capture temporal dependencies over long horizons prior works either improve framewise features with transformer or refine framewise predictions with learned action features. However they are compu…

2024

Learning to Segment Referred Objects from Narrated Egocentric Videos

CVPR 2024poster

Egocentric videos provide a first-person perspective of the wearer's activities involving simultaneous interactions with multiple objects. In this work we propose the task of weakly-supervised Narration-based Video Object Segmentation (NVOS). Given an egocentric video clip and a narration of the wea…

Cited by 6SourcePDFScholar
2022

Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

CVPR 2022poster

Open-vocabulary instance segmentation aims at segmenting novel classes without mask annotations. It is an important step toward reducing laborious human supervision. Most existing works first pretrain a model on captioned images covering many novel classes and then finetune it on limited base classe…

Cited by 104PDFcodeScholar
2021

Interaction Compass: Multi-Label Zero-Shot Learning of Human-Object Interactions via Spatial Relations

ICCV 2021poster

We study the problem of multi-label zero-shot recognition in which labels are in the form of human-object interactions (combinations of actions on objects), each image may contain multiple interactions and some interactions do not have training images. We propose a novel compositional learning frame…

Cited by 10PDFcodeScholar
2021

Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence Alignment

CVPR 2021poster

We address the problem of unsupervised localization of key-steps and feature learning in instructional videos using both visual and language instructions. Our key observation is that the sequences of visual and linguistic key-steps are weakly aligned: there is an ordered one-to-one correspondence be…

Cited by 50PDFcodeScholar
2021

Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces Learning

ICCV 2021poster

We address the problem of learning to segment actions from weakly-annotated videos, i.e., videos accompanied by transcripts (ordered list of actions). We propose a framework in which we model actions with a union of low-dimensional subspaces, learn the subspaces using transcripts and refine video fe…

Cited by 28PDFcodeScholar
2019

Deep Supervised Summarization: Algorithm and Application to Learning Instructions

NeurIPS 2019poster

We address the problem of finding representative points of datasets by learning from multiple datasets and their ground-truth summaries. We develop a supervised subset selection framework, based on the facility location utility function, which learns to map datasets to their ground-truth representat…

Cited by 9SourcePDFScholar