← Search

Fida Mohammad Thoker

5 accepted papers

2026

TrackMAE: Video Representation Learning via Track Mask and Predict

CVPR 2026

Masked video modeling (MVM) has emerged as a simple and scalable self-supervised pretraining paradigm, but only encodes motion information implicitly, limiting the encoding of temporal dynamics in the learned representations. As a result, such models struggle on motion-centric tasks that require fin

Cited by 0SourcecodeScholar
2025

SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning

CVPR 2025poster

Masked video modeling, such as VideoMAE, is an effective paradigm for video self-supervised learning (SSL). However, they are primarily based on reconstructing pixel level details on natural videos which have substantial temporal redundancy, limiting their capability for semantic representation and…

2024

SIGMA: Sinkhorn-Guided Masked Video Modeling

ECCV 2024poster

"Video-based pretraining offers immense potential for learning strong visual representations on an unprecedented scale. Recently, masked video modeling methods have shown promising scalability, yet fall short in capturing higher-level semantics due to reconstructing predefined low-level targets such…

Cited by 2SourcePDFScholar
2023

Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization

ICCV 2023poster

We propose a self-supervised method for learning motion-focused video representations. Existing approaches minimize distances between temporally augmented videos, which maintain high spatial similarity. We instead propose to learn similarities between videos with identical local motion dynamics but…

Cited by 13PDFcodeScholar
2022

How Severe Is Benchmark-Sensitivity in Video Self-Supervised Learning?

ECCV 2022poster

"Despite the recent success of video self-supervised learning models, there is much still to be understood about their generalization capability. In this paper, we investigate how sensitive video self-supervised learning is to the current conventional benchmark and whether methods generalize beyond…