← Search

Shubham Tulsiani

63 accepted papers

2026

Contact-guided Real2Sim from Monocular Video with Planar Scene Primitives

ICLR 2026poster

We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human--scene reconstruction relies on data-driven priors and joint optimization with no physics in the loop, or recovers noisy geometry with artifacts that cause motion-tr…

Cited by 0SourcecodeScholar
2026

DemoDiffusion: One-Shot Human Imitation Using Pre-Trained Diffusion Policy

ICRA 2026poster

We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring task-specific training or paired human-robot data. Our approach is based on two insights. First, the hand motion in a human demonstration provides …

2026

E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

CVPR 2026

Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations from multi-view images. In this paper, we present E-RayZer, a self-supervised 3D vision model that learns geometrically

Cited by 0SourcecodeScholar
2026

EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing

CVPR 2026

High-fidelity generative video editing has seen significant quality improvements by leveraging pre-trained video foundation models. However, their computational cost is a major bottleneck, as they are often designed to inefficiently process the full video context regardless of the inpainting mask's

Cited by 0SourcecodeScholar
2026

Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning

CVPR 2026

Current feed-forward 3D/4D reconstruction systems rely on dense geometry and pose supervision - expensive to obtain at scale and particularly scarce for dynamic real-world scenes. We present Flow3r, a framework that augments visual geometry learning with dense 2D correspondences ('flow') as supervis

Cited by 0SourcecodeScholar
2026

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

RSS 2026poster

We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii…

Cited by 0SourceScholar
2026

Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models

ICML 2026poster

We present a mechanism to steer the sampling diversity of denoising diffusion and flow matching models, allowing users to sample from a sharper or broader distribution than the training distribution. We build on the observation that these models leverage (learned) score functions of noisy data distr…

Cited by 0SourceScholar
2025

AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis

CVPR 2025poster

We explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image pairs. Our hypothesis is that the lack of high-quality, co-regis…

Cited by 1SourcePDFScholar
2025

DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion

CVPR 2025poster

Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning approach that directly infers 3D scene geometry and camera poses…

2025

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

CoRL 2025poster

How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in terms of predicting motion information from web data through human video generation and conditioning a robot policy on the generated video. Instead of…

Cited by 0SourceScholar
2025

LightSwitch: Multi-view Relighting with Material-guided Diffusion

ICCV 2025poster

Recent approaches for 3D relighting have shown promise in integrating 2D image relighting generative priors to alter the appearance of a 3D representation while preserving the underlying structure. Nevertheless, generative priors used for 2D relighting that directly relight from an input image do no…

Cited by 0SourcePDFScholar
2025

SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation

CVPR 2025poster

We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortless editing. SceneFactor enables text-guided 3D scene synthesis through our factored diffusion formulation, leveraging latent semantic and geometric manifolds for ge…

Cited by 6SourcePDFScholar
2025

UniPhy: Learning a Unified Constitutive Model for Inverse Physics Simulation

CVPR 2025poster

We propose UniPhy, a common latent-conditioned neural constitutive model that can encode the physical properties of diverse materials. At inference UniPhy allows `inverse simulation' i.e. inferring material properties by optimizing the scene-specific latent to match the available observations via di…

Cited by 0SourcePDFScholar
2024

Cameras as Rays: Pose Estimation via Ray Diffusion

ICLR 2024oral

Estimating camera poses is a fundamental task for 3D reconstruction and remains challenging given sparsely sampled views (<10). In contrast to existing approaches that pursue top-down prediction of global parametrizations of camera extrinsics, we propose a distributed representation of camera pose t…

Cited by 60SourcePDFScholar
2024

G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis

CVPR 2024poster

We propose G-HOP a denoising diffusion based generative prior for hand-object interactions that allows modeling both the 3D object and a human hand conditioned on the object category. To learn a 3D spatial diffusion model that can capture this joint distribution we represent the human hand via a ske…

Cited by 0SourcePDFScholar
2024

MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation

CVPR 2024poster

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models these generations are not 3D-consistent and require a distillation process to generate…

Cited by 19SourcePDFScholar
2024

RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking

ICRA 2024poster

The grand aim of having a single robot that can manipulate arbitrary objects in diverse settings is at odds with the paucity of robotics datasets. Acquiring and growing such datasets is strenuous due to manual efforts, operational costs, and safety challenges. A path toward such a universal agent re…

Cited by 131SourcecodeScholar
2024

Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis

NeurIPS 2024poster

Inferring the 3D structure underlying a set of multi-view images typically requires solving two co-dependent tasks -- accurate 3D reconstruction requires precise camera poses, and predicting camera poses relies on (implicitly or explicitly) modeling the underlying 3D. The classical framework of anal…

Cited by 1SourcePDFScholar
2024

Towards Generalizable Zero-Shot Manipulation via Translating Human Interaction Plans

ICRA 2024poster

We pursue the goal of developing robots that can interact zero-shot with generic unseen objects via a diverse repertoire of manipulation skills and show how passive human videos can serve as a rich source of data for learning such generalist robots. Unlike typical robot learning approaches which dir…

Cited by 44SourcecodeScholar
2024

Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

ECCV 2024poster

"We seek to learn a generalizable goal-conditioned policy that enables diverse robot manipulation — interacting with unseen objects in novel scenes without test-time adaptation. While typical approaches rely on a large amount of demonstration data for such generalization, we propose an approach that…

2024

UpFusion: Novel View Diffusion from Unposed Sparse View Observations

ECCV 2024poster

"We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for generic objects given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically rely on camera poses to geometrically aggregate info…

2023

Affordance Diffusion: Synthesizing Hand-Object Interactions

CVPR 2023poster

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting objects into a user-specified region. In contrast, in this work…

2023

Analogy-Forming Transformers for Few-Shot 3D Parsing

ICLR 2023poster

We present Analogical Networks, a model that segments 3D object scenes with analogical reasoning: instead of mapping a scene to part segments directly, our model first retrieves related scenes from memory and their corresponding part structures, and then predicts analogous part structures in the inp…

Cited by 5SourcePDFScholar
2023

Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction Clips

ICCV 2023oral

We tackle the task of reconstructing hand-object interactions from short video clips. Given an input video, our approach casts 3D inference as a per-video optimization and recovers a neural 3D representation of the object shape, as well as the time-varying motion and hand articulation. While the inp…

Cited by 47PDFScholar
2023

Manipulate by Seeing: Creating Manipulation Controllers from Pre-Trained Representations

ICCV 2023oral

The field of visual representation learning has seen explosive growth in the past years, but its benefits in robotics have been surprisingly limited so far. Prior work uses generic visual representations as a basis to learn (task-specific) robot action policies (e.g., via behavior cloning). While th…

Cited by 18PDFcodeScholar
2023

SparseFusion: Distilling View-Conditioned Diffusion for 3D Reconstruction

CVPR 2023poster

We propose SparseFusion, a sparse view 3D reconstruction approach that unifies recent advances in neural rendering and probabilistic image generation. Existing approaches typically build on neural rendering with re-projected features but fail to generate unseen regions or handle uncertainty under la…

2022

AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation

CVPR 2022poster

Powerful priors allow us to perform inference with insufficient information. In this paper, we propose an autoregressive prior for 3D shapes to solve multimodal 3D tasks such as shape completion, reconstruction, and generation. We model the distribution over 3D shapes as a non-sequential autoregress…

Cited by 268PDFcodeScholar
2022

Monocular Dynamic View Synthesis: A Reality Check

NeurIPS 2022accept

We study the recent progress on dynamic view synthesis (DVS) from monocular video. Though existing approaches have demonstrated impressive results, we show a discrepancy between the practical capture process and the existing experimental protocols, which effectively leaks in multi-view signals durin…

2022

Pre-Train, Self-Train, Distill: A Simple Recipe for Supersizing 3D Reconstruction

CVPR 2022poster

Our work learns a unified model for single-view 3D reconstruction of objects from hundreds of semantic categories. As a scalable alternative to direct 3D supervision, our work relies on segmented image collections for learning 3D of generic categories. Unlike prior works that use similar supervision…

Cited by 44PDFScholar
2022

RelPose: Predicting Probabilistic Relative Rotation for Single Objects in the Wild

ECCV 2022poster

"We describe a data-driven method for inferring the camera viewpoints given multiple images of an arbitrary object. This task is a core component of classic geometric pipelines such as SfM and SLAM, and also serves as a vital pre-processing requirement for contemporary neural approaches (e.g. NeRF)…

Cited by 94SourcePDFScholar
2021

A Differentiable Recipe for Learning Visual Non-Prehensile Planar Manipulation

CoRL 2021poster

Specifying tasks with videos is a powerful technique towards acquiring novel and general robot skills. However, reasoning over mechanics and dexterous interactions can make it challenging to scale visual learning for contact-rich manipulation. In this work, we focus on the problem of visual dexterou…

Cited by 4SourcecodeScholar
2021

NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild

NeurIPS 2021poster

Recent history has seen a tremendous growth of work exploring implicit representations of geometry and radiance, popularized through Neural Radiance Fields (NeRF). Such works are fundamentally based on a (implicit) {\em volumetric} representation of occupancy, allowing them to model diverse scene s…

2021

No RL, No Simulation: Learning to Navigate without Navigating

NeurIPS 2021poster

Most prior methods for learning navigation policies require access to simulation environments, as they need online policy interaction and rely on ground-truth maps for rewards. However, building simulators is expensive (requires manual effort for each and every scene) and creates challenges in trans…

Cited by 93SourcePDFScholar
2021

Where2Act: From Pixels to Actions for Articulated 3D Objects

ICCV 2021poster

One of the fundamental goals of visual perception is to allow agents to meaningfully interact with their environment. In this paper, we take a step towards that long-term goal -- we extract highly localized actionable information related to elementary actions such as pushing or pulling for articulat…

Cited by 202PDFcodeScholar
2020

Efficient Bimanual Manipulation Using Learned Task Schemas

ICRA 2020poster

We address the problem of effectively composing skills to solve sparse-reward tasks in the real world. Given a set of parameterized skills (such as exerting a force or doing a top grasp at a location), our goal is to learn policies that invoke these skills to efficiently solve such tasks. Our insigh…

Cited by 84SourceScholar
2020

Intrinsic Motivation for Encouraging Synergistic Behavior

ICLR 2020poster

We study the role of intrinsic motivation as an exploration bias for reinforcement learning in sparse-reward synergistic tasks, which are tasks where multiple agents must work together to achieve a goal they could not individually. Our key idea is that a good guiding principle for intrinsic motivati…

Cited by 32SourceScholar
2020

See, Hear, Explore: Curiosity via Audio-Visual Association

NeurIPS 2020poster

Exploration is one of the core challenges in reinforcement learning. A common formulation of curiosity-driven exploration uses the difference between the real future and the future predicted by a learned model. However, predicting the future is an inherently difficult task which can be ill-posed in…

2020

Use the Force, Luke! Learning to Predict Physical Forces by Simulating Effects

CVPR 2020oral

When we humans look at a video of human-object interaction, we can not only infer what is happening but we can even extract actionable information and imitate those interactions. On the other hand, current recognition or geometric approaches lack the physicality of action representation. In this pap…

Cited by 57PDFcodeScholar
2019

3D-RelNet: Joint Object and Relational Network for 3D Prediction

ICCV 2019poster

We propose an approach to predict the 3D shape and pose for the objects present in a scene. Existing learning based methods that pursue this goal make independent predictions per object, and do not leverage the relationships amongst them. We argue that reasoning about these relationships is crucial,…

Cited by 58PDFScholar
2019

Object-centric Forward Modeling for Model Predictive Control

CoRL 2019

We present an approach to learn an object-centric forward model, and show that this allows us to plan for sequences of actions to achieve distant desired goals. We propose to model a scene as a collection of objects, each with an explicit spatial location and implicit visual feature, and learn to mo

2019

Order-Aware Generative Modeling Using the 3D-Craft Dataset

ICCV 2019poster

In this paper, we study the problem of sequentially building houses in the game of Minecraft, and demonstrate that learning the ordering can make for more effective autoregressive models. Given a partially built house made by a human player, our system tries to place additional blocks in a human-lik…

Cited by 9PDFcodeScholar
2018

Factoring Shape, Pose, and Layout From the 2D Image of a 3D Scene

CVPR 2018poster

The goal of this paper is to take a single 2D image of a scene and recover the 3D structure in terms of a small set of factors: a layout representing the enclosing surfaces as well as a set of objects represented in terms of shape and pose. We propose a convolutional neural network-based approach to…

Cited by 157SourcePDFScholar
2018

Learning Category-Specific Mesh Reconstruction from Image Collections

ECCV 2018poster

We present a learning framework for recovering the 3D shape, camera, and texture of an object from a single image. The shape is represented as a deformable 3D mesh model of an object category where a shape is parameterized by a learned mean shape and per-instance predicted deformation. Our approach…

2018

Multi-View Consistency as Supervisory Signal for Learning Shape and Pose Prediction

CVPR 2018poster

We present a framework for learning single-view shape and pose prediction without using direct supervision for either. Our approach allows leveraging multi-view observations from unknown poses as supervisory signal during training. Our proposed training setup enforces geometric consistency between t…

Cited by 232SourcePDFScholar
2017

Learning Shape Abstractions by Assembling Volumetric Primitives

CVPR 2017poster

We present a learning framework for abstracting complex shapes by learning to assemble objects using 3D volumetric primitives. In addition to generating simple and geometrically interpretable explanations of 3D objects, our framework also allows us to automatically discover and exploit consistent st…

Cited by 402PDFcodeScholar
2017

Multi-View Supervision for Single-View Reconstruction via Differentiable Ray Consistency

CVPR 2017oral

We study the notion of consistency between a 3D shape and a 2D observation and propose a differentiable formulation which allows computing gradients of the 3D shape given an observation from an arbitrary view. We do so by reformulating view consistency using a differentiable ray consistency (DRC) te…

Cited by 646PDFScholar
2015

Category-Specific Object Reconstruction From a Single Image

CVPR 2015poster

Object reconstruction from a single image -- in the wild -- is a problem where we can make progress and get meaningful results today. This is the main message of this paper, which introduces an automated pipeline with pixels as inputs and 3D surfaces of various rigid categories as outputs in images…

Cited by 421SourcePDFScholar