← Search

Pedro Morgado

19 accepted papers

2025

From Prototypes to General Distributions: An Efficient Curriculum for Masked Image Modeling

CVPR 2025poster

Masked Image Modeling (MIM) has emerged as a powerful self-supervised learning paradigm for visual representation learning, enabling models to acquire rich visual representations by predicting masked portions of images from their visible regions. While this approach has shown promising results, we h…

Cited by 0SourcePDFScholar
2025

TrackVerse: A Large-Scale Object-Centric Video Dataset for Image-Level Representation Learning

ICCV 2025accepted

Video data inherently captures rich, dynamic contexts that reveal objects in varying poses, interactions, and state transitions, offering rich potential for unsupervised object representation learning. However, most prior representation learning methods rely on static image datasets like ImageNet, w…

2024

Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling

CVPR 2024poster

Humans possess a remarkable ability to integrate auditory and visual information enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues demonstrated through cognitive psychology and neuroscience research offers promising potential for developing mu…

2023

A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition

ICML 2023poster

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independently for each task. However, given the interconnected nature of source localizat…

2023

Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?

ICCV 2023poster

Vision-language models such as CLIP learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through few-shot prompt tuning. We find that such prompt tuning process is highly robust to label noises. This intrigues us to…

Cited by 19PDFcodeScholar
2022

Learning State-Aware Visual Representations from Audible Interactions

NeurIPS 2022accept

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In result, several large egocentric datasets of interaction-rich…

2020

Learning Representations from Audio-Visual Spatial Alignment

NeurIPS 2020poster

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual correspondence (AVC) predict whether audio and video clips origi…

2020

Solving Long-tailed Recognition with Deep Realistic Taxonomic Classifier

ECCV 2020poster

Long-tail recognition tackles the natural non-uniformly distributed data in real-world scenarios. While modern classifiers perform well on populated classes, its performance degrades significantly on tail classes. Humans, however, are less affected by this since, when confronted with uncertain examp…

2018

Self-Supervised Generation of Spatial Audio for 360° Video

NeurIPS 2018poster

We introduce an approach to convert mono audio recorded by a 360° video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360° video viewing, but spatial audio microphones are still rare in curr…

Cited by 201SourcePDFScholar