2026
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
CVPR 2026
Transformers are remarkably versatile, suggesting the existence of generic inductive biases beneficial across modalities. In this work, we explore a new way to instil such biases in vision transformers (ViTs) through pretraining on procedurally generated data devoid of visual or semantic content. We