← Search

Saurabh Saxena

11 accepted papers

2026

ORBIT: Benchmarking SfM in the Wild with 360deg Video

CVPR 2026

Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes.Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard t

Cited by 0SourceScholar
2025

Controlling Space and Time with Diffusion Models

ICLR 2025poster

We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with…

2025

High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion

AAAI 2025technical

Despite the recent progress, existing frame interpolation methods still struggle with processing extremely high resolution input and handling challenging cases such as repetitive textures, thin objects, and large motion. To address these issues, we introduce a patch-based cascaded pixel diffusion mo…

Cited by 0SourcePDFScholar
2025

RoMo: Robust Motion Segmentation Improves Structure from Motion

ICCV 2025poster

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. Estimating accurate camera poses from videos through structure-from-motion (SfM) relies on robustly separating static and dynamic parts of a video. We propose a novel approach…

Cited by 0SourcePDFScholar
2024

NeRFiller: Completing Scenes via Generative 3D Inpainting

CVPR 2024poster

We propose NeRFiller an approach that completes missing portions of a 3D capture via generative 3D inpainting using off-the-shelf 2D visual generative models. Often parts of a captured 3D scene or object are missing due to mesh reconstruction failures or a lack of observations (e.g. contact regions…

Cited by 33SourcePDFScholar
2023

A Generalist Framework for Panoptic Segmentation of Images and Videos

ICCV 2023poster

Panoptic segmentation assigns semantic and instance ID labels to every pixel of an image. As permutations of instance IDs are also valid solutions, the task requires learning of high-dimensional one-to-many mapping. As a result, state-of-the-art approaches use customized architectures and task-speci…

Cited by 123PDFcodeScholar
2023

The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth Estimation

NeurIPS 2023oral

Denoising diffusion probabilistic models have transformed image generation with their impressive fidelity and diversity. We show that they also excel in estimating optical flow and monocular depth, surprisingly without task-specific architectures and loss functions that are predominant for these tas…

Cited by 96SourcePDFScholar
2022

A Unified Sequence Interface for Vision Tasks

NeurIPS 2022accept

While language tasks are naturally expressed in a single, unified, modeling framework, i.e., generating sequences of tokens, this has not been the case in computer vision. As a result, there is a proliferation of distinct architectures and loss functions for different vision tasks. In this work we s…

2022

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

NeurIPS 2022accept

We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image gene…

Cited by 6404SourcePDFScholar
2022

Pix2seq: A Language Modeling Framework for Object Detection

ICLR 2022poster

We present Pix2Seq, a simple and generic framework for object detection. Unlike existing approaches that explicitly integrate prior knowledge about the task, we cast object detection as a language modeling task conditioned on the observed pixel inputs. Object descriptions (e.g., bounding boxes and c…

2017

Large-Scale Evolution of Image Classifiers

ICML 2017poster

Neural networks have proven effective at solving difficult problems but designing their architectures can be challenging, even for image classification problems alone. Our goal is to minimize human participation, so we employ evolutionary algorithms to discover such networks automatically. Despite s…

Cited by 2148SourcePDFScholar