← Search

Ali Diba

11 accepted papers

2021

3D CNNs With Adaptive Temporal Feature Resolutions

CVPR 2021poster

While state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there…

Cited by 39PDFcodeScholar
2021

Temporally-Weighted Hierarchical Clustering for Unsupervised Action Segmentation

CVPR 2021poster

Action segmentation refers to inferring boundaries of semantically consistent visual concepts in videos and is an important requirement for many video understanding tasks. For this and other video understanding tasks, supervised approaches have achieved encouraging performance but require a high vol…

Cited by 81PDFcodeScholar
2021

Vi2CLR: Video and Image for Visual Contrastive Learning of Representation

ICCV 2021poster

In this paper, we introduce a novel self-supervised visual representation learning method which understands both images and videos in a joint learning fashion. The proposed neural network architecture and objectives are designed to obtain two different Convolutional Neural Networks for solving visua…

Cited by 66PDFScholar
2020

Large Scale Holistic Video Understanding

ECCV 2020poster

Video recognition has been advanced in recent years by benchmarks with rich annotations. However, research is still mainly limited to human action or sports recognition - focusing on a highly specific video understanding task and thus leaving a significant gap towards describing the overall content…

2018

Classification-Driven Dynamic Image Enhancement

CVPR 2018poster

Convolutional neural networks rely on image texture and structure to serve as discriminative features to classify the image content. Image enhancement techniques can be used as preprocessing steps to help improve the overall image quality and in turn improve the overall effectiveness of a CNN. Exis…

Cited by 88SourcePDFScholar
2018

Spatio-Temporal Channel Correlation Networks for Action Classification

ECCV 2018poster

The work in this paper is driven by the question if spatio-temporal correlations are enough for 3D convolutional neural networks (CNN)? Most of the traditional 3D networks use local spatio-temporal features. We introduce a new block that models correlations between channels of a 3D CNN with respect…

Cited by 239SourcePDFScholar
2016

DeepCAMP: Deep Convolutional Action & Attribute Mid-Level Patterns

CVPR 2016poster

The recognition of human actions and the determination of human attributes are two tasks that call for fine-grained classification. Indeed, often rather small and inconspicuous objects and features have to be detected to tell their classes apart. In order to deal with this challenge, we propose a…

Cited by 58PDFScholar
2015

DeepProposal: Hunting Objects by Cascading Deep Convolutional Layers

ICCV 2015poster

In this paper we evaluate the quality of the activation layers of a convolutional neural network (CNN) for the generation of object proposals. We generate hypotheses in a sliding-window fashion over different activation layers and show that the final convolutional layers can find the object of inter…

Cited by 150PDFScholar