← Search

Anand Bhattad

18 accepted papers

2026

Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning

ICRA 2026poster

We study view-invariant imitation learning by explicitly conditioning policies on camera extrinsics. Using Plücker embeddings of per-pixel rays, we show that conditioning on extrinsics significantly improves generalization across viewpoints for standard behavior cloning policies, including ACT, Diff…

2026

Generalizable Sparse-View 3D Reconstruction from Unconstrained Images

CVPR 2026

Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization with appearance embeddings or dynamic masks, requiring extensive per-scene training and failin

Cited by 0SourcecodeScholar
2026

Generative Blocks World: Moving Things Around in Pictures

ICLR 2026poster

We describe Generative Blocks World to interact with the scene of a generated image by manipulating simple geometric abstractions. Our method represents scenes as assemblies of convex 3D primitives, and the same scene can be represented by different numbers of primitives, allowing an editor to move…

Cited by 0SourceScholar
2026

Stronger Semantic Encoders Can Harm Relighting Performance: A Probe of Visual Priors via Augmented Latent Intrinsics

ICML 2026poster

Image-to-image relighting requires representations that disentangle scene properties from illumination. Recent methods rely on latent intrinsic representations but remain under-constrained and often fail on challenging materials such as metal and glass. A natural hypothesis is that stronger pretrain…

Cited by 0SourceScholar
2025

LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting

CVPR 2025poster

We introduce LumiNet, a novel architecture that leverages generative models and latent intrinsic representations for transferring lighting from one image to another. Given a source image and a target lighting image, LumiNet generates a relit version of the source scene that captures the target's lig…

Cited by 3SourcePDFScholar
2025

ScribbleLight: Single Image Indoor Relighting with Scribbles

CVPR 2025poster

Image-based relighting of indoor rooms creates an immersive virtual understanding of the space, which is useful for interior design, virtual staging, and real estate. Relighting indoor rooms from a single image is especially challenging due to complex illumination interactions between multiple light…

Cited by 2SourcePDFScholar
2025

Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting

NeurIPS 2025poster

This paper proposes a novel scene understanding task called Visual Jenga. Drawing inspiration from the game Jenga, the proposed task involves progressively removing objects from a single image until only the background remains. Just as Jenga players must understand structural dependencies to maintai…

Cited by 0SourceScholar
2024

From an Image to a Scene: Learning to Imagine the World from a Million 360° Videos

NeurIPS 2024poster

Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training mod…

2024

Latent Intrinsics Emerge from Training to Relight

NeurIPS 2024spotlight

Image relighting is the task of showing what a scene from a source image would look like if illuminated differently. Inverse graphic schemes recover an explicit representation of geometry and a set of chosen intrinsics, then relight with some form of renderer. But error control for inverse graphic…

Cited by 1SourcePDFScholar
2024

Shadows Don't Lie and Lines Can't Bend! Generative Models don't know Projective Geometry...for now

CVPR 2024poster

Generative models can produce impressively realistic images. This paper demonstrates that generated images have geometric features different from those of real images. We build a set of collections of generated images prequalified to fool simple signal-based classifiers into believing they are real.…

2024

Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion

ECCV 2024poster

"We introduce , a training-free video editing algorithm for localized semantic edits. allows users to use any editing software, including Photoshop and generative inpainting, to modify the first frame; it automatically propagates those changes, with semantic, spatial, and temporally consistent motio…

Cited by 7SourcePDFScholar
2023

Improving Equivariance in State-of-the-Art Supervised Depth and Normal Predictors

ICCV 2023poster

Dense depth and surface normal predictors should possess the equivariant property to cropping-and-resizing -- cropping the input image should result in cropping the same output image. However, we find that state-of-the-art depth and normal predictors, despite having strong performances, surprisingly…

Cited by 1PDFcodeScholar
2023

OBJECT 3DIT: Language-guided 3D-aware Image Editing

NeurIPS 2023poster

Existing image editing tools, while powerful, typically disregard the underlying 3D geometry from which the image is projected. As a result, edits made using these tools may become detached from the geometry and lighting conditions that are at the foundation of the image formation process; such edit…

Cited by 37SourcePDFScholar
2022

DIVeR: Real-Time and Accurate Neural Radiance Fields With Deterministic Integration for Volume Rendering

CVPR 2022oral

DIVeR builds on the key ideas of NeRF and its variants -- density models and volume rendering -- to learn 3D object models that can be rendered realistically from small numbers of images. In contrast to all previous NeRF methods, DIVeR uses deterministic rather than stochastic estimates of the volum…

Cited by 85PDFcodeScholar
2021

View Generalization for Single Image Textured 3D Models

CVPR 2021poster

Humans can easily infer the underlying 3D geometry and texture of an object only from a single 2D image. Current computer vision methods can do this, too, but suffer from view generalization problems -- the models inferred tend to make poor predictions of appearance in novel views. As for generaliza…

Cited by 35PDFScholar
2020

Unrestricted Adversarial Examples via Semantic Manipulation

ICLR 2020poster

Machine learning models, especially deep neural networks (DNNs), have been shown to be vulnerable against adversarial examples which are carefully crafted samples with a small magnitude of the perturbation. Such adversarial perturbations are usually restricted by bounding their $\mathcal{L}_p$ norm…

Cited by 178SourceScholar