← Search

Mohammadreza Zolfaghari

4 accepted papers

2021

CrossCLR: Cross-Modal Contrastive Learning for Multi-Modal Video Representations

ICCV 2021poster

Contrastive learning allows us to flexibly define powerful losses by contrasting positive pairs from sets of negative samples. Recently, the principle has also been used to learn cross-modal embeddings for video and text, yet without exploiting its full potential. In particular, previous losses do n…

Cited by 172PDFScholar
2020

COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning

NeurIPS 2020poster

Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the…

2018

ECO: Efficient Convolutional Network for Online Video Understanding

ECCV 2018poster

The state of the art in video understanding suffers from two problems: (1) The major part of reasoning is performed locally in the video, thus missing important relationships within actions that span several seconds. (2) While there are local methods with fast per-frame processing, the processing of…

2017

Chained Multi-Stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection

ICCV 2017poster

General human action recognition requires understanding of various visual cues. In this paper, we propose a network architecture that computes and integrates the most important visual cues for action recognition: pose, motion, and the raw images. For the integration, we introduce a Markov chain mode…

Cited by 277PDFScholar