← Search

Rahul Ravishankar

2 accepted papers

2025

An Empirical Study of Autoregressive Pre-training from Videos

ICCV 2025poster

We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a…

Cited by 0SourcePDFScholar
2025

Scaling Properties of Diffusion Models For Perceptual Tasks

CVPR 2025poster

In this paper, we argue that iterative computation with diffusion models offers a powerful paradigm for not only generation but also visual perception tasks. We unify tasks such as depth estimation, optical flow, and amodal segmentation under the framework of image-to-image translation, and show how…

Cited by 4SourcePDFScholar