← Search

Kalyan Sunkavalli

49 accepted papers

2026

E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

CVPR 2026

Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations from multi-view images. In this paper, we present E-RayZer, a self-supervised 3D vision model that learns geometrically

Cited by 0SourcecodeScholar
2026

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

CVPR 2026

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability. Our framework efficiently compresses multiple image observat

Cited by 0SourcecodeScholar
2025

4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time

NeurIPS 2025poster

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction model that takes input from unconstrained views and timesta…

Cited by 0SourceScholar
2025

RayZer: A Self-supervised Large View Synthesis Model

ICCV 2025poster

We present RayZer, a self-supervised multi-view 3D Vision model trained without any 3D supervision, i.e., camera poses and scene geometry, while exhibiting emerging 3D awareness. Concretely, RayZer takes unposed and uncalibrated images as input, recovers camera parameters, reconstructs a scene repre…

Cited by 0SourcePDFScholar
2024

DATENeRF: Depth-Aware Text-based Editing of NeRFs

ECCV 2024poster

"Recent diffusion models have demonstrated impressive capabilities for text-based 2D image editing. Applying similar ideas to edit a NeRF scene [?] remains challenging as editing 2D frames individually does not produce multiview-consistent results. We make the key observation that the geometry of a…

Cited by 4SourcePDFScholar
2024

DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction Model

ICLR 2024spotlight

We propose DMV3D, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model incorporates a triplane NeRF representation and, functioning as a denoiser, can denoise noisy multi-view images via 3D NeRF reconstru…

2024

GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting

ECCV 2024poster

"We propose , a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in ∼0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture; we patchify input posed images, pass the concatenated multi-view…

2024

Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction Model

ICLR 2024poster

Text-to-3D with diffusion models has achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from slow inference, low diversity and Janus problems, or are feed-forward methods that generate low-quality results due to…

Cited by 250SourcePDFScholar
2024

LRM: Large Reconstruction Model for Single Image to 3D

ICLR 2024oral

We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable t…

Cited by 411SourcePDFScholar
2024

LightIt: Illumination Modeling and Control for Diffusion Models

CVPR 2024poster

We introduce LightIt a method for explicit illumination control for image generation. Recent generative methods lack lighting control which is crucial to numerous artistic aspects of image generation such as setting the overall mood or cinematic appearance. To overcome these limitations we propose t…

Cited by 16SourcePDFScholar
2024

Neural Directional Encoding for Efficient and Accurate View-Dependent Appearance Modeling

CVPR 2024highlight

Novel-view synthesis of specular objects like shiny metals or glossy paints remains a significant challenge. Not only the glossy appearance but also global illumination effects including reflections of other objects in the environment are critical components to faithfully reproduce a scene. In this…

2024

PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction

ICLR 2024spotlight

We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in ~1.3 seconds on a single A100 GPU. PF-LRM is a highly scalable method utilizing self-atten…

2023

Interactive Portrait Harmonization

ICLR 2023poster

Current image harmonization methods consider the entire background as the guidance for harmonization. However, this may limit the capability for user to choose any specific object/person in the background to guide the harmonization. To enable flexible interaction between user and harmonization, we i…

Cited by 8SourcePDFScholar
2023

PaletteNeRF: Palette-Based Appearance Editing of Neural Radiance Fields

CVPR 2023poster

Recent advances in neural radiance fields have enabled the high-fidelity 3D reconstruction of complex scenes for novel view synthesis. However, it remains underexplored how the appearance of such representations can be efficiently edited while maintaining photorealism. In this work, we present Palet…

Cited by 63SourcePDFScholar
2022

NeRFusion: Fusing Radiance Fields for Large-Scale Scene Reconstruction

CVPR 2022oral

While NeRF has shown great success for neural reconstruction and rendering, its limited MLP capacity and long per-scene optimization times make it challenging to model large-scale indoor scenes. In contrast, classical 3D reconstruction methods can handle large-scale scenes but do not produce realist…

Cited by 125PDFcodeScholar
2022

PhotoScene: Photorealistic Material and Lighting Transfer for Indoor Scenes

CVPR 2022poster

Most indoor 3D scene reconstruction methods focus on recovering 3D geometry and scene layout. In this work, we go beyond this to propose PhotoScene, a framework that takes input image(s) of a scene along with approximately aligned CAD geometry (either reconstructed automatically or manually specifie…

Cited by 31PDFcodeScholar
2022

Physically-Based Editing of Indoor Scene Lighting from a Single Image

ECCV 2022poster

"We present a method to edit complex indoor lighting from a single image with its predicted depth and light source segmentation masks. This is an extremely challenging problem that requires modeling complex light transport, and disentangling HDR lighting from material and geometry with only a partia…

Cited by 61SourcePDFScholar
2022

Point-NeRF: Point-Based Neural Radiance Fields

CVPR 2022oral

Volumetric neural rendering methods like NeRF generate high-quality view synthesis results but are optimized per-scene leading to prohibitive reconstruction time. On the other hand, deep multi-view stereo methods can quickly reconstruct scene geometry via direct network inference. Point-NeRF combine…

Cited by 701PDFcodeScholar
2021

Deep Denoising of Flash and No-Flash Pairs for Photography in Low-Light Environments

CVPR 2021poster

We introduce a neural network-based method to denoise pairs of images taken in quick succession in low-light environments, with and without a flash. Our goal is to produce a high-quality rendering of the scene that preserves the color and mood from the ambient illumination of the noisy no-flash imag…

Cited by 25PDFScholar
2021

NeuTex: Neural Texture Mapping for Volumetric Neural Rendering

CVPR 2021poster

Recent work has demonstrated that volumetric scene representations combined with differentiable volume rendering can enable photo-realistic rendering for challenging scenes that mesh reconstruction fails on. However, these methods entangle geometry and appearance in a ""black-box"" volume that canno…

Cited by 116PDFScholar
2021

OpenRooms: An Open Framework for Photorealistic Indoor Scene Datasets

CVPR 2021poster

We propose a novel framework for creating large-scale photorealistic datasets of indoor scenes, with ground truth geometry, material, lighting and semantics. Our goal is to make the dataset creation process widely accessible, allowing researchers to transform scans into datasets with highquality gro…

Cited by 93PDFScholar
2021

SSH: A Self-Supervised Framework for Image Harmonization

ICCV 2021poster

Image harmonization aims to improve the quality of image compositing by matching the "appearance"" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead,…

Cited by 96PDFcodeScholar
2020

Basis Prediction Networks for Effective Burst Denoising With Large Kernels

CVPR 2020poster

Bursts of images exhibit significant self-similarity across both time and space. This motivates a representation of the kernels as linear combinations of a small set of basis elements. To this end, we introduce a novel basis prediction network that, given an input burst, predicts a set of global bas…

Cited by 86PDFScholar
2020

Deep 3D Capture: Geometry and Reflectance From Sparse Multi-View Images

CVPR 2020poster

We introduce a novel learning-based method to reconstruct the high-quality geometry and complex, spatially-varying BRDF of an arbitrary object from a sparse set of only six images captured by wide-baseline cameras under collocated point lighting. We first estimate per-view depth maps using a deep mu…

Cited by 97PDFScholar
2020

Deep Reflectance Volumes: Relightable Reconstructions from Multi-View Photometric Images

ECCV 2020poster

We present a deep learning approach to reconstruct scene appearance from unstructured images captured under collocated point lighting. At the heart of Deep Reflectance Volumes is a novel volumetric scene representation consisting of opacity, surface normal and reflectance voxel grids. We present a n…

Cited by 133SourcePDFScholar
2020

Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single Image

CVPR 2020oral

We propose a deep inverse rendering framework for indoor scenes. From a single RGB image of an arbitrary indoor scene, we obtain a complete scene reconstruction, estimating shape, spatially-varying lighting, and spatially-varying, non-Lambertian surface reflectance. Our novel inverse rendering netwo…

Cited by 294PDFcodeScholar
2020

Single View Metrology in the Wild

ECCV 2020poster

Most 3D reconstruction methods may only recover scene properties up to a global scale ambiguity. We present a novel approach to single view metrology that can recover the absolute scale of a scene represented by 3D heights of objects or camera height above the ground as well as camera parameters of…

2019

All-Weather Deep Outdoor Lighting Estimation

CVPR 2019poster

We present a neural network that predicts HDR outdoor illumination from a single LDR image. At the heart of our work is a method to accurately learn HDR lighting from LDR panoramas under any weather condition. We achieve this by training another CNN (on a combination of synthetic and real images) to…

Cited by 88PDFScholar
2019

Deep CG2Real: Synthetic-to-Real Translation via Image Disentanglement

ICCV 2019poster

We present a method to improve the visual realism of low-quality, synthetic images, e.g. OpenGL renderings. Training an unpaired synthetic-to-real translation network in image space is severely under-constrained and produces visible artifacts. Instead, we propose a semi-supervised approach that oper…

Cited by 44PDFScholar
2019

Deep Parametric Indoor Lighting Estimation

ICCV 2019poster

We present a method to estimate lighting from a single image of an indoor scene. Previous work has used an environment map representation that does not account for the localized nature of indoor lighting. Instead, we represent lighting as a set of discrete 3D lights with geometric and photometric pa…

Cited by 159PDFScholar
2019

Fast Spatially-Varying Indoor Lighting Estimation

CVPR 2019oral

We propose a real-time method to estimate spatially-varying indoor lighting from a single RGB image. Given an image and a 2D location in that image, our CNN estimates a 5th order spherical harmonic representation of the lighting at the given location in less than 20ms on a laptop mobile graphics car…

Cited by 167PDFScholar
2019

Learning to Separate Multiple Illuminants in a Single Image

CVPR 2019poster

We present a method to separate a single image captured under two illuminants, with different spectra, into the two images corresponding to the appearance of the scene under each individual illuminant. We do this by training a deep neural network to predict the per-pixel reflectance chromaticity of…

Cited by 18PDFScholar
2018

A Perceptual Measure for Deep Single Image Camera Calibration

CVPR 2018poster

Most current single image camera calibration methods rely on specific image features or user input, and cannot be applied to natural images captured in uncontrolled settings. We propose inferring directly camera calibration parameters from a single image using a deep convolutional neural network. Th…

Cited by 143SourcePDFScholar
2018

Compositing-aware Image Search

ECCV 2018poster

We present a new image search technique that, given a background image, returns compatible foreground objects for image compositing tasks. The compatibility of a foreground object and a background scene depends on various aspects such as semantics, surrounding context, geometry, style and color. How…

Cited by 21SourcePDFScholar
2018

Fast Video Object Segmentation by Reference-Guided Mask Propagation

CVPR 2018poster

We present an efficient method for the semi-supervised video object segmentation. Our method achieves accuracy competitive with state-of-the-art methods while running in a fraction of time compared to others. To this end, we propose a deep Siamese encoder-decoder network that is designed to take adv…

Cited by 512SourcePDFScholar
2018

Illuminant Spectra-Based Source Separation Using Flash Photography

CVPR 2018poster

Real-world lighting often consists of multiple illuminants with different spectra. Separating and manipulating these illuminants in post-process is a challenging problem that requires either significant manual input or calibrated scene geometry and lighting. In this work, we leverage a flash/no-flas…

Cited by 16SourcePDFScholar
2018

MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics

ECCV 2018poster

Long-term human motion can be represented as a series of motion modes—motion sequences that capture short-term temporal dynamics—with transitions between them. We leverage this structure and present a novel Motion Transformation Variational Auto-Encoders (MT-VAE) for learning motion sequence generat…

Cited by 183SourcePDFScholar
2018

Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image

ECCV 2018poster

We propose a material acquisition system that can recover the spatially-varying BRDF and normal map of a near-planar surface from a single image captured by a handheld mobile phone camera. Our technique images the surface under arbitrary environment lighting with the flash turned on, thereby avoidin…

Cited by 184SourcePDFScholar
2017

Deep Outdoor Illumination Estimation

CVPR 2017oral

We present a CNN-based technique to estimate high-dynamic range outdoor illumination from a single low dynamic range image. To train the CNN, we leverage a large dataset of outdoor panoramas. We fit a low-dimensional physically-based outdoor illumination model to the skies in these panoramas giving…

Cited by 272PDFScholar
2017

Neural Face Editing With Intrinsic Image Disentangling

CVPR 2017oral

Traditional face editing methods often require a number of sophisticated and task specific algorithms to be applied one after the other --- a process that is tedious, fragile, and computationally intensive. In this paper, we propose an end-to-end generative adversarial network that infers a face-spe…

Cited by 340PDFcodeScholar
2017

Reflectance Capture Using Univariate Sampling of BRDFs

ICCV 2017poster

We propose the use of a light-weight setup consisting of a collocated camera and light source --- commonly found on mobile devices --- to reconstruct surface normals and spatially-varying BRDFs of near-planar material samples. A collocated setup provides only a 1-D "univariate" sampling of the 4-D B…

Cited by 73PDFScholar
2017

Scene Parsing With Global Context Embedding

ICCV 2017poster

We present a scene parsing method that utilizes global context information based on both the parametric and non-parametric models. Compared to previous methods that only exploit the local relationship between objects, we train a context network based on scene similarities to generate feature represe…

Cited by 70PDFcodeScholar
2016

Automatic Content-Aware Color and Tone Stylization

CVPR 2016spotlight

We introduce a new technique that automatically generates diverse, visually compelling stylizations for a photograph in an unsupervised manner. We achieve this by learning style ranking for a given input using a large photo collection and selecting a diverse subset of matching styles for final style…

Cited by 92PDFScholar
2015

PatchMatch-Based Automatic Lattice Detection for Near-Regular Textures

ICCV 2015poster

In this work, we investigate the problem of automatically inferring the lattice structure of near-regular textures (NRT) in real-world images. Our technique leverages the PatchMatch algorithm for finding k-nearest-neighbor (kNN) correspondences in an image. We use these kNNs to recover an initial es…

Cited by 28PDFScholar