← Search

Humam Alwassel

3 accepted papers

2020

Self-Supervised Learning by Cross-Modal Audio-Video Clustering

NeurIPS 2020spotlight

Visual and audio modalities are highly correlated, yet they contain different information. Their strong correlation makes it possible to predict the semantics of one from the other with good accuracy. Their intrinsic differences make cross-modal prediction a potentially more rewarding pretext task f…

2018

Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization

ECCV 2018poster

State-of-the-art temporal action detectors inefficiently search the entire video for specific actions. Despite the encouraging progress these methods achieve, it is crucial to design automated approaches that only explore parts of the video which are the most relevant to the actions being searched f…

2018

Diagnosing Error in Temporal Action Detectors

ECCV 2018poster

Despite the recent progress in video understanding and the continuous rate of improvement in temporal action localization throughout the years, it is still unclear how far (or close?) we are to solving the problem. To this end, we introduce a new diagnostic tool to analyze the performance of tempora…