CVPR 20260 citations

In Pursuit of Pixel Supervision for Visual Pre-training

Lihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei, Dong Wang, Abdelrahman Mohamed, Saining Xie, Hengshuang Zhao

Abstract

Data matters. In computer vision, data (or pixels) are the primary source of information containing signals that span from low-level attributes to high-level concepts. At scale, the success of modern vision systems has been closely tied to how data is curated for semantic understanding (e.g., ImageNet). Recent trends in spatial intelligence and physical world understanding further highlight the importance of real-world signals that preserve spatial structure, beyond purely semantic signals. This motivates a shift toward curating data that better captures spatially grounded information across diverse environments. In this work, we demonstrate that training on 2B web-crawled images with a self-curation strategy on masked autoencoder (MAE) can learn strong representations for dense prediction tasks, while remaining simple, stable, and efficient. Our model, codenamed "Pixio", is an enhanced masked autoencoder (MAE) with more challenging pre-training tasks and more capable architectures. Pixio yields dense representations achieving promising results across a wide range of dense prediction tasks in the wild, including monocular depth estimation (e.g., Depth Anything), feed-forward 3D reconstruction (i.e., MapAnything), and visual segmentation (e.g., SAM). Our results suggest that data curation can significantly contribute to dense representation learning.

BibTeX
@inproceedings{cvpr2026_inpursuitofpixel,
  title = {In Pursuit of Pixel Supervision for Visual Pre-training},
  author = {Lihe Yang and Shang-Wen Li and Yang Li and Xinjie Lei and Dong Wang and Abdelrahman Mohamed and Saining Xie and Hengshuang Zhao and Kaiming He and Hu Xu},
  booktitle = {CVPR 2026},
  year = {2026}
}
In Pursuit of Pixel Supervision for Visual Pre-training · CVPR 2026