← Search

Varun Jampani

106 accepted papers

2026

GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra

ICLR 2026poster

Monocular 3D reconstruction methods and vision-language models (VLMs) demonstrate impressive results on standard benchmarks, yet their true understanding of geometric properties remains unclear. We introduce GIQ, a comprehensive benchmark specifically designed to evaluate the geometric reasoning cap…

Cited by 0SourcecodeScholar
2026

ReLi3D: Relightable Multi-view 3D Reconstruction with Disentangled Illumination

ICLR 2026poster

Reconstructing 3D assets from images has long required separate pipelines for geometry reconstruction, material estimation, and illumination recovery, each with distinct limitations and computational overhead. We present MIDR-3D, the first unified end-to-end pipeline that simultaneously reconstructs…

Cited by 0SourcecodeScholar
2025

CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussians

CVPR 2025poster

Recent breakthroughs in text-guided image generation have significantly advanced the field of 3D generation. While generating a single high-quality 3D object is now feasible, generating multiple objects with reasonable interactions within a 3D space, a.k.a. compositional 3D generation, presents subs…

Cited by 4SourcePDFScholar
2025

FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image

ICCV 2025poster

We present a novel framework for generating high-quality, animatable 4D avatar from a single image. While recent advances have shown promising results in 4D avatar creation, existing methods either require extensive multiview data or struggle with geometry accuracy and identity consistency. To addre…

Cited by 0SourcePDFScholar
2025

HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Models

ICCV 2025poster

We introduce HouseCrafter, a novel approach that can lift a 2D floorplan into a complete large 3D indoor scene (e.g., a house). Our key insight is to adapt a 2D diffusion model, which is trained on web-scale images, to generate consistent multi-view color (RGB) and depth (D) images across different…

Cited by 0SourcePDFScholar
2025

MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

ICLR 2025spotlight

Estimating geometry from dynamic scenes, where objects move and deform over time, remains a core challenge in computer vision. Current approaches often rely on multi-stage pipelines or global optimizations that decompose the problem into subtasks, like depth and flow, leading to complex systems pron…

2025

Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models

ICCV 2025poster

In this paper, we analyze the viewpoint stability of foundational models - specifically, their sensitivity to changes in viewpoint- and define instability as significant feature variations resulting from minor changes in viewing angle, leading to generalization gaps in 3D reasoning tasks. We investi…

Cited by 0SourcePDFScholar
2025

PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling

ICCV 2025poster

Skinning and rigging are fundamental components in animation, articulated object reconstruction, motion transfer, and 4D generation. Existing approaches predominantly rely on Linear Blend Skinning (LBS), due to its simplicity and differentiability. However, LBS introduces artifacts such as volume lo…

Cited by 0SourcePDFScholar
2025

SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement

CVPR 2025poster

We present SF3D, a novel method for rapid and high-quality textured object mesh reconstruction from a single image in just 0.5 seconds. Unlike most existing approaches, SF3D is explicitly trained for mesh generation, incorporating a fast UV unwrapping technique that enables swift texture generation…

Cited by 27SourcePDFScholar
2025

SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

CVPR 2025poster

We study the problem of single-image 3D object reconstruction. Recent works have diverged into two directions: regression-based modeling and generative modeling. Regression methods efficiently infer visible surfaces, but struggle with occluded regions. Generative methods handle uncertain regions bet…

2025

SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation

ICCV 2025poster

We present Stable Video 4D 2.0 (SV4D 2.0), a multi-view video diffusion model for dynamic 3D asset generation. Compared to its predecessor SV4D, SV4D 2.0 is more robust to occlusions and large motion, generalizes better to real-world videos, and produces higher-quality outputs in terms of detail sha…

Cited by 0SourcePDFScholar
2025

SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency

ICLR 2025poster

We present Stable Video 4D (SV4D) — a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to…

2025

SViM3D: Stable Video Material Diffusion for Single Image 3D Generation

ICCV 2025poster

We present Stable Video Materials 3D (SViM3D), a framework to predict multi-view consistent physically based rendering (PBR) materials, given a single image. Recently, video diffusion models have been successfully used to reconstruct 3D objects from a single image efficiently. However, reflectance i…

Cited by 0SourcePDFScholar
2025

Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation

NeurIPS 2025poster

Recent advances in video generation have enabled high-fidelity video synthesis from user provided prompts. However, existing models and benchmarks fail to capture the complexity and requirements of professional video generation. Towards that goal, we introduce Stable Cinemetrics, a structured evalua…

Cited by 0SourceScholar
2025

Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation

NeurIPS 2025spotlight

We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to produce kinematic parts --- structural components aligned with…

Cited by 0SourceScholar
2025

Stable Virtual Camera: Generative View Synthesis with Diffusion Models

ICCV 2025poster

We present \underline \text S tabl\underline \text e \underline \text V irtual C\underline \text a mera (Seva), a generalist diffusion model that creates novel views of a scene, given any number of input views and target cameras.Existing works struggle to generate either large viewpoint changes…

Cited by 0SourcePDFScholar
2025

SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D Editing

AAAI 2025technical

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits across multiple viewpoints remains a challenge. While the it…

Cited by 0SourcePDFScholar
2025

Unified Dense Prediction of Video Diffusion

CVPR 2025poster

We present a unified network for simultaneously generating videos and their corresponding entity segmentation and depth maps from text prompts. We utilize colormap to represent entity masks and depth maps, tightly integrating dense prediction with RGB video generation. Introducing dense prediction i…

Cited by 0SourcePDFScholar
2024

3D Congealing: 3D-Aware Image Alignment in the Wild

ECCV 2024poster

"We propose , a novel problem of 3D-aware alignment for 2D images capturing semantically similar objects. Given a collection of unlabeled Internet images, our goal is to associate the shared semantic parts from the inputs and aggregate the knowledge from 2D images to a shared 3D canonical space. We…

Cited by 0SourcePDFScholar
2024

Alchemist: Parametric Control of Material Properties with Diffusion Models

CVPR 2024poster

We propose a method to control material attributes of objects like roughness metallic albedo and transparency in real images. Our method capitalizes on the generative prior of text-to-image models known for photorealism employing a scalar value and instructions to alter low-level material properties…

Cited by 20SourcePDFScholar
2024

ConDense: Consistent 2D-3D Pre-training for Dense and Sparse Features from Multi-View Images

ECCV 2024oral

"To advance the state of the art in the creation of 3D foundation models, this paper introduces the framework for 3D pre-training utilizing existing pre-trained 2D networks and large-scale multi-view datasets. We propose a novel 2D-3D joint training scheme to extract co-embedded 2D and 3D features i…

Cited by 5SourcePDFScholar
2024

DiffusionLight: Light Probes for Free by Painting a Chrome Ball

CVPR 2024poster

We present a simple yet effective technique to estimate lighting in a single input image. Current techniques rely heavily on HDR panorama datasets to train neural networks to regress an input with limited field-of-view to a full environment map. However these approaches often struggle with real-worl…

2024

HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models

CVPR 2024poster

Personalization has emerged as a prominent aspect within the field of generative AI enabling the synthesis of individuals in diverse contexts and styles while retaining high-fidelity to their identities. However the process of personalization presents inherent challenges in terms of time and memory…

Cited by 191SourcePDFScholar
2024

Learning Action and Reasoning-Centric Image Editing from Videos and Simulation

NeurIPS 2024spotlight

An image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current *general* instruction-guided editing models have significant shortcomings with action and r…

Cited by 4SourcePDFScholar
2024

MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation

CVPR 2024poster

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models these generations are not 3D-consistent and require a distillation process to generate…

Cited by 19SourcePDFScholar
2024

NeRFiller: Completing Scenes via Generative 3D Inpainting

CVPR 2024poster

We propose NeRFiller an approach that completes missing portions of a 3D capture via generative 3D inpainting using off-the-shelf 2D visual generative models. Often parts of a captured 3D scene or object are missing due to mesh reconstruction failures or a lack of observations (e.g. contact regions…

Cited by 33SourcePDFScholar
2024

OmniControl: Control Any Joint at Any Time for Human Motion Generation

ICLR 2024poster

We present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis trajectory, OmniControl can incorporate flexible spatial con…

2024

One-Shot Open Affordance Learning with Foundation Models

CVPR 2024poster

We introduce One-shot Open Affordance Learning (OOAL) where a model is trained with just one example per base object category but is expected to identify novel objects and affordances. While vision-language models excel at recognizing novel objects and scenes they often struggle to understand finer…

2024

PRISM: A New Lens for Improved Color Understanding

EMNLP 2024industry

While image-text pre-trained models, such as CLIP, have demonstrated impressive capabilities in learning robust text and image representations, a critical area for substantial improvement remains—precise color understanding. In this paper, we address this limitation by introducing PRISM, a simple ye…

Cited by 0SourcePDFScholar
2024

Probing the 3D Awareness of Visual Foundation Models

CVPR 2024poster

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task their intermediate representations are useful for other visual tasks such as detection and segmentation. Given…

2024

SHINOBI: Shape and Illumination using Neural Object Decomposition via BRDF Optimization In-the-wild

CVPR 2024poster

We present SHINOBI an end-to-end framework for the reconstruction of shape material and illumination from object images captured with varying lighting pose and background. Inverse rendering of an object based on unconstrained image collections is a long-standing challenge in computer vision and grap…

Cited by 6SourcePDFScholar
2024

Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence

CVPR 2024poster

While pre-trained large-scale vision models have shown significant promise for semantic correspondence their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation o…

2024

UniGS: Unified Representation for Image Generation and Segmentation

CVPR 2024poster

This paper introduces a novel unified representation of diffusion models for image generation and segmentation. Specifically we use a colormap to represent entity-level masks addressing the challenge of varying entity numbers while aligning the representation closely with the image RGB domain. Two n…

2024

WordRobe: Text-Guided Generation of Textured 3D Garments

ECCV 2024poster

"In this paper, we tackle a new and challenging problem of text-driven generation of 3D garments with high-quality textures. We propose, WordRobe, a novel framework for the generation of unposed & textured 3D garment meshes from user-friendly text prompts. We achieve this by first learning a latent…

Cited by 10SourcePDFScholar
2024

ZeroShape: Regression-based Zero-shot Shape Reconstruction

CVPR 2024poster

We study the problem of single-image zero-shot 3D shape reconstruction. Recent works learn zero-shot shape reconstruction through generative modeling of 3D assets but these models are computationally expensive at train and inference time. In contrast the traditional approach to this problem is regre…

2023

A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence

NeurIPS 2023poster

Text-to-image diffusion models have made significant advances in generating and editing high-quality images. As a result, numerous approaches have explored the ability of diffusion model features to understand and process single images for downstream tasks, e.g., classification, semantic segmentati…

2023

ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections

NeurIPS 2023poster

Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wi…

Cited by 23SourcePDFScholar
2023

ASIC: Aligning Sparse in-the-wild Image Collections

ICCV 2023oral

We present a method for joint alignment of sparse in-the-wild image collections of an object category. Most prior works assume either ground-truth keypoint annotations or a large dataset of images of a single object category. However, neither of the above assumptions hold true for the long-tail of t…

Cited by 21PDFcodeScholar
2023

DreamBooth3D: Subject-Driven Text-to-3D Generation

ICCV 2023poster

We present DreamBooth3D, an approach to personalize text-to-3D generative models from as few as 3-6 casually captured images of a subject. Our approach combines recent advances in personalizing text-to-image models (DreamBooth) with text-to-3D generation (DreamFusion). We find that naively combining…

Cited by 211PDFScholar
2023

DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

CVPR 2023poster

Large text-to-image models achieved a remarkable leap in the evolution of AI, enabling high-quality and diverse synthesis of images from a given text prompt. However, these models lack the ability to mimic the appearance of subjects in a given reference set and synthesize novel renditions of them in…

2023

Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery From Sparse Image Ensemble

CVPR 2023poster

Automatically estimating 3D skeleton, shape, camera viewpoints, and part articulation from sparse in-the-wild image ensembles is a severely under-constrained and challenging problem. Most prior methods rely on large-scale image datasets, dense temporal correspondence, or human annotations like camer…

2023

KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models

ACL 2023industry

Image ad understanding is a crucial task with wide real-world applications. Although highly challenging with the involvement of diverse atypical scenes, real-world entities, and reasoning over scene-texts, how to interpret image ads is relatively under-explored, especially in the era of foundational…

Cited by 5SourcePDFScholar
2023

LOCATE: Localize and Transfer Object Parts for Weakly Supervised Affordance Grounding

CVPR 2023poster

Humans excel at acquiring knowledge through observation. For example, we can learn to use new tools by watching demonstrations. This skill is fundamental for intelligent systems to interact with the world. A key step to acquire this skill is to identify what part of the object affords each action, w…

Cited by 51SourcePDFScholar
2023

LU-NeRF: Scene and Pose Estimation by Synchronizing Local Unposed NeRFs

ICCV 2023poster

A critical obstacle preventing NeRF models from being deployed broadly in the wild is their reliance on accurate camera poses. Consequently, there is growing interest in extending NeRF models to jointly optimize camera poses and scene representation, which offers an alternative to off-the-shelf SfM…

Cited by 34PDFScholar
2023

LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

NeurIPS 2023poster

Attaining a high degree of user controllability in visual generation often requires intricate, fine-grained inputs like layouts. However, such inputs impose a substantial burden on users when compared to simple text inputs. To address the issue, we study how Large Language Models (LLMs) can serve as…

2023

MetaCLUE: Towards Comprehensive Visual Metaphors Research

CVPR 2023poster

Creativity is an indispensable part of human cognition and also an inherent part of how we make sense of the world. Metaphorical abstraction is fundamental in communicating creative ideas through nuanced relationships between abstract concepts such as feelings. While computer vision benchmarks and a…

2023

NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations

NeurIPS 2023poster

Recent advances in neural reconstruction enable high-quality 3D object reconstruction from casually captured image collections. Current techniques mostly analyze their progress on relatively simple image collections where SfM techniques can provide ground-truth (GT) camera poses. We note that SfM te…

2023

NoisyTwins: Class-Consistent and Diverse Image Generation Through StyleGANs

CVPR 2023poster

StyleGANs are at the forefront of controllable image generation as they produce a latent space that is semantically disentangled, making it suitable for image editing and manipulation. However, the performance of StyleGANs severely degrades when trained via class-conditioning on large-scale long-tai…

2023

ShapeClipper: Scalable 3D Shape Learning From Single-View Images via Geometric and CLIP-Based Consistency

CVPR 2023poster

We present ShapeClipper, a novel method that reconstructs 3D object shapes from real-world single-view RGB images. Instead of relying on laborious 3D, multi-view or camera pose annotation, ShapeClipper learns shape reconstruction from a set of single-view segmented images. The key idea is to facilit…

Cited by 22SourcePDFScholar
2023

Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

ICLR 2023poster

Large-scale diffusion models have achieved state-of-the-art results on text-to-image synthesis (T2I) tasks. Despite their ability to generate high-quality yet creative images, we observe that attribution-binding and compositional capabilities are still considered major challenging issues, especially…

2022

Amplitude Spectrum Transformation for Open Compound Domain Adaptive Semantic Segmentation

AAAI 2022technical

Open compound domain adaptation (OCDA) has emerged as a practical adaptation setting which considers a single labeled source domain against a compound of multi-modal unlabeled target data in order to generalize better on novel unseen domains. We hypothesize that an improved disentanglement of domain…

Cited by 12SourcePDFScholar
2022

Balancing Discriminability and Transferability for Source-Free Domain Adaptation

ICML 2022spotlight

Conventional domain adaptation (DA) techniques aim to improve domain transferability by learning domain-invariant representations; while concurrently preserving the task-discriminability knowledge gathered from the labeled source data. However, the requirement of simultaneous access to labeled sourc…

2022

CPL: Counterfactual Prompt Learning for Vision and Language Models

EMNLP 2022main

Prompt tuning is a new few-shot transfer learning technique that only tunes the learnable prompt for pre-trained vision and language models such as CLIP. However, existing prompt tuning methods tend to learn spurious or entangled representations, which leads to poor generalization to unseen concepts…

Cited by 27SourcePDFScholar
2022

Concurrent Subsidiary Supervision for Unsupervised Source-Free Domain Adaptation

ECCV 2022poster

"The prime challenge in unsupervised domain adaptation (DA) is to mitigate the domain shift between the source and target domains. Prior DA works show that pretext tasks could be used to mitigate this domain shift by learning domain invariant representations. However, in practice, we find that most…

2022

Hierarchical Semantic Regularization of Latent Spaces in StyleGANs

ECCV 2022poster

"Progress in GANs has enabled the generation of high-resolution photorealistic images of astonishing quality. StyleGANs allow for compelling attribute modification on such images via mathematical operations on the latent style vectors in the W/W+ space that effectively modulate the rich hierarchical…

Cited by 10SourcePDFScholar
2022

Improving GANs for Long-Tailed Data through Group Spectral Regularization

ECCV 2022poster

"Deep long-tailed learning aims to train useful deep networks on practical, real-world imbalanced distributions, wherein most labels of the tail classes are associated with a few samples. There has been a large body of work to train discriminative models for visual recognition on long-tailed distrib…

2022

LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part Discovery

NeurIPS 2022accept

Creating high-quality articulated 3D models of animals is challenging either via manual creation or using 3D scanning tools. Therefore, techniques to reconstruct articulated 3D objects from 2D images are crucial and highly useful. In this work, we propose a practical problem setting to estimate 3D…

2022

Learning ABCs: Approximate Bijective Correspondence for Isolating Factors of Variation With Weak Supervision

CVPR 2022oral

Representational learning forms the backbone of most deep learning applications, and the value of a learned representation is intimately tied to its information content regarding different factors of variation. Finding good representations depends on the nature of supervision and the learning algori…

Cited by 2PDFcodeScholar
2022

Planes vs. Chairs: Category-Guided 3D Shape Learning without Any 3D Cues

ECCV 2022poster

"We present a novel 3D shape reconstruction method which learns to predict an implicit 3D shape representation from a single RGB image. Our approach uses a set of single-view images of multiple object categories without viewpoint annotation, forcing the model to learn across multiple object categori…

Cited by 16SourcePDFScholar
2022

Polynomial Neural Fields for Subband Decomposition and Manipulation

NeurIPS 2022accept

Neural fields have emerged as a new paradigm for representing signals, thanks to their ability to do it compactly while being easy to optimize. In most applications, however, neural fields are treated like a black box, which precludes many signal manipulation tasks. In this paper, we propose a new c…

2022

SAMURAI: Shape And Material from Unconstrained Real-world Arbitrary Image collections

NeurIPS 2022accept

Inverse rendering of an object under entirely unknown capture conditions is a fundamental challenge in computer vision and graphics. Neural approaches such as NeRF have achieved photorealistic results on novel view synthesis, but they require known camera poses. Solving this problem with unknown cam…

Cited by 81SourcePDFScholar
2022

SOMSI: Spherical Novel View Synthesis With Soft Occlusion Multi-Sphere Images

CVPR 2022poster

Spherical novel view synthesis (SNVS) is the task of estimating 360 views at dynamic novel views given a set of 360 input views. Prior arts learn multi-sphere image (MSI) representations that enables fast rendering times but are only limited to modelling low-dimensional color values. Modelling high-…

Cited by 17PDFcodeScholar
2022

Subsidiary Prototype Alignment for Universal Domain Adaptation

NeurIPS 2022accept

Universal Domain Adaptation (UniDA) deals with the problem of knowledge transfer between two datasets with domain-shift as well as category-shift. The goal is to categorize unlabeled target samples, either into one of the "known" categories or into a single "unknown" category. A major problem in Uni…

Cited by 25SourcePDFScholar
2022

Uncertainty-Aware Adaptation for Self-Supervised 3D Human Pose Estimation

CVPR 2022poster

The advances in monocular 3D human pose estimation are dominated by supervised techniques that require large-scale 2D/3D pose annotations. Such methods often behave erratically in the absence of any provision to discard unfamiliar out-of-distribution data. To this end, we cast the 3D human pose lear…

Cited by 45PDFScholar
2022

ViDT: An Efficient and Effective Fully Transformer-based Object Detector

ICLR 2022poster

Transformers are transforming the landscape of computer vision, especially for recognition tasks. Detection transformers are the first fully end-to-end learning systems for object detection, while vision transformers are the first fully transformer-based architecture for image classification. In thi…

2021

Adaptive Prototype Learning and Allocation for Few-Shot Segmentation

CVPR 2021poster

Prototype learning is extensively used for few-shot segmentation. Typically, a single prototype is obtained from the support feature by averaging the global object information. However, using one prototype to represent all the information may lead to ambiguities. In this paper, we propose two novel…

Cited by 473PDFcodeScholar
2021

Aligning Silhouette Topology for Self-Adaptive 3D Human Pose Recovery

NeurIPS 2021poster

Articulation-centric 2D/3D pose supervision forms the core training objective in most existing 3D human pose estimation techniques. Except for synthetic source environments, acquiring such rich supervision for each real target domain at deployment is highly inconvenient. However, we realize that sta…

Cited by 11SourcePDFScholar
2021

AutoFlow: Learning a Better Training Set for Optical Flow

CVPR 2021poster

Synthetic datasets play a critical role in pre-training CNN models for optical flow, but they are painstaking to generate and hard to adapt to new applications. To automate the process, we present AutoFlow, a simple and effective method to render training data for optical flow that optimizes the per…

Cited by 132PDFcodeScholar
2021

Generalize Then Adapt: Source-Free Domain Adaptive Semantic Segmentation

ICCV 2021poster

Unsupervised domain adaptation (DA) has gained substantial interest in semantic segmentation. However, almost all prior arts assume concurrent access to both labeled source and unlabeled target, making them unsuitable for scenarios demanding source-free adaptation. In this work, we enable source-fre…

Cited by 142PDFcodeScholar
2021

Implicit-PDF: Non-Parametric Representation of Probability Distributions on the Rotation Manifold

ICML 2021spotlight

In the deep learning era, the vast majority of methods to predict pose from a single image are trained to classify or regress to a single given ground truth pose per image. Such methods have two main shortcomings, i) they cannot represent uncertainty about the predictions, and ii) they cannot handle…

2021

Infinite Nature: Perpetual View Generation of Natural Scenes From a Single Image

ICCV 2021poster

We introduce the problem of perpetual view generation - long-range generation of novel views corresponding to an arbitrarily long camera trajectory given a single image. This is a challenging problem that goes far beyond the capabilities of current view synthesis methods, which quickly degenerate wh…

Cited by 169PDFcodeScholar
2021

LASR: Learning Articulated Shape Reconstruction From a Monocular Video

CVPR 2021poster

Remarkable progress has been made in 3D reconstruction of rigid structures from a video or a collection of images. However, it is still challenging to reconstruct nonrigid structures from RGB inputs, due to the under-constrained nature of this problem. While template-based approaches, such as parame…

Cited by 129PDFcodeScholar
2021

NeRD: Neural Reflectance Decomposition From Image Collections

ICCV 2021poster

Decomposing a scene into its shape, reflectance, and illumination is a challenging but important problem in computer vision and graphics. This problem is inherently more challenging when the illumination is not a single light source under laboratory conditions but is instead an unconstrained environ…

Cited by 547PDFcodeScholar
2021

Neural-PIL: Neural Pre-Integrated Lighting for Reflectance Decomposition

NeurIPS 2021poster

Decomposing a scene into its shape, reflectance and illumination is a fundamental problem in computer vision and graphics. Neural approaches such as NeRF have achieved remarkable success in view synthesis, but do not explicitly perform decomposition and instead operate exclusively on radiance (the p…

2021

Non-local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation

NeurIPS 2021poster

Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target environment is highly inconvenient. To this end, we cast 3D pose lear…

Cited by 12SourcePDFScholar
2021

Robust Visual Reasoning via Language Guided Neural Module Networks

NeurIPS 2021poster

Neural module networks (NMN) are a popular approach for solving multi-modal tasks such as visual question answering (VQA) and visual referring expression recognition (REF). A key limitation in prior implementations of NMN is that the neural modules do not effectively capture the association between…

Cited by 27SourcePDFScholar
2021

SLIDE: Single Image 3D Photography With Soft Layering and Depth-Aware Inpainting

ICCV 2021poster

Single image 3D photography enables viewers to view a still image from novel viewpoints. Recent approaches combine monocular depth networks with inpainting networks to achieve compelling results. A drawback of these techniques is the use of hard depth layering, making them unable to model intricate…

Cited by 51PDFcodeScholar
2021

ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction

NeurIPS 2021spotlight

We introduce ViSER, a method for recovering articulated 3D shapes and dense3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple camera views, strong category-specific priors, or 2D keypoint supervision. We show that no…

2020

Appearance Consensus Driven Self-Supervised Human Mesh Recovery

ECCV 2020poster

We present a self-supervised human mesh recovery framework to infer human pose and shape from monocular images in the absence of any paired supervision. Recent advances have shifted the interest towards directly regressing parameters of a parametric human model by supervising them on large-scale, im…

Cited by 48SourcePDFScholar
2020

DeepGMR: Learning Latent Gaussian Mixture Models for Registration

ECCV 2020poster

Point cloud registration is a fundamental problem in 3D computer vision, graphics and robotics. For the last few decades, existing registration algorithms have struggled in situations with large transformations, noise, and time constraints. In this paper, we introduce Deep Gaussian Mixture Registrat…

2020

From Image Collections to Point Clouds With Self-Supervised Shape and Pose Networks

CVPR 2020poster

Reconstructing 3D models from 2D images is one of the fundamental problems in computer vision. In this work, we propose a deep learning technique for 3D object reconstruction from a single image. Contrary to recent works that either use 3D supervision or multi-view supervision, we use only single vi…

Cited by 40PDFcodeScholar
2020

Generative View Synthesis: From Single-view Semantics to Novel-view Images

NeurIPS 2020poster

Content creation, central to applications such as virtual reality, can be tedious and time-consuming. Recent image synthesis methods simplify this task by offering tools to generate new views from as little as a single input image, or by converting a semantic map into a photorealistic image. We pro…

2020

Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis

CVPR 2020oral

Camera captured human pose is an outcome of several sources of variation. Performance of supervised 3D pose estimation approaches comes at the cost of dispensing with variations, such as shape and appearance, that may be useful for solving other related tasks. As a result, the learned model not only…

Cited by 109PDFScholar
2020

Self-Supervised Viewpoint Learning From Image Collections

CVPR 2020poster

Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively easy to mine many unlabeled images of an object category from t…

Cited by 45PDFcodeScholar
2020

Self-supervised Single-view 3D Reconstruction via Semantic Consistency

ECCV 2020poster

We learn a self-supervised, single-view 3D reconstruction model that predicts the 3D mesh shape, texture and camera pose of a target object with a collection of 2D images and silhouettes. The proposed method does not necessitate 3D supervision, manually annotated keypoints, multi-view images of an o…

Cited by 200SourcePDFScholar
2020

Two-Shot Spatially-Varying BRDF and Shape Estimation

CVPR 2020poster

Capturing the shape and spatially-varying appearance (SVBRDF) of an object from images is a challenging task that has applications in both computer vision and graphics. Traditional optimization-based approaches often need a large number of images taken from multiple views in a controlled environment…

Cited by 107PDFcodeScholar
2019

Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation

CVPR 2019poster

We address the unsupervised learning of several interconnected problems in low-level vision: single view depth prediction, camera motion estimation, optical flow, and segmentation of a video into the static scene and moving regions. Our key insight is that these four fundamental vision problems are…

Cited by 742PDFcodeScholar
2019

Pixel-Adaptive Convolutional Neural Networks

CVPR 2019poster

Convolutions are the fundamental building blocks of CNNs. The fact that their weights are spatially shared is one of the main reasons for their widespread use, but it is also a major limitation, as it makes convolutions content-agnostic. We propose a pixel-adaptive convolution (PAC) operation, a sim…

Cited by 383PDFcodeScholar
2019

SCOPS: Self-Supervised Co-Part Segmentation

CVPR 2019poster

Parts provide a good intermediate representation of objects that is robust with respect to camera, pose and appearance variations. Existing work on part segmentation is dominated by supervised approaches that rely on large amounts of manual annotations and also can not generalize to unseen object ca…

Cited by 178PDFScholar
2019

SENSE: A Shared Encoder Network for Scene-Flow Estimation

ICCV 2019oral

We introduce a compact network for holistic scene flow estimation, called SENSE, which shares common encoder features among four closely-related tasks: optical flow estimation, disparity estimation from stereo, occlusion estimation, and semantic segmentation. Our key insight is that sharing features…

Cited by 88PDFcodeScholar
2018

Learning Superpixels With Segmentation-Aware Affinity Loss

CVPR 2018poster

Superpixel segmentation has been widely used in many computer vision tasks. Existing superpixel algorithms are mainly based on hand-crafted features, which often fail to preserve weak object boundaries. In this work, we leverage deep neural networks to facilitate extracting superpixels from images.…

Cited by 151SourcePDFScholar
2018

SPLATNet: Sparse Lattice Networks for Point Cloud Processing

CVPR 2018poster

We present a network architecture for processing point clouds that directly operates on a collection of points represented as a sparse set of samples in a high-dimensional lattice. Naively applying convolutions on this lattice scales poorly, both in terms of memory and computational cost, as the siz…

2018

Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video Interpolation

CVPR 2018poster

Given two consecutive frames, video interpolation aims at generating intermediate frame(s) to form both spatially and temporally coherent video sequences. While most existing methods focus on single-frame interpolation, we propose an end-to-end convolutional neural network for variable-length multi-…

Cited by 981SourcePDFScholar
2018

Switchable Temporal Propagation Network

ECCV 2018poster

Videos contain highly redundant information between frames. Such redundancy has been studied extensively in video compression and encoding but is less explored for more advanced video processing. In this paper, we propose a learnable unified framework for propagating a variety of visual properties o…

Cited by 49SourcePDFScholar
2016

Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs and Bilateral Neural Networks

CVPR 2016poster

Bilateral filters have wide spread use due to their edge-preserving properties. The common use case is to manually choose a parametric filter type, usually a Gaussian filter. In this paper, we will generalize the parametrization and in particular derive a gradient descent algorithm so the filter par…

Cited by 244PDFScholar
2016

Optical Flow With Semantic Segmentation and Localized Layers

CVPR 2016spotlight

Existing optical flow methods make generic, spatially homogeneous, assumptions about the spatial structure of the flow. In reality, optical flow varies across an image depending on object class. Simply put, different objects move differently. Here we exploit recent advances in static semantic scene…

Cited by 251PDFScholar
2015

Consensus Message Passing for Layered Graphical Models

AISTATS 2015poster

Generative models provide a powerful framework for probabilistic reasoning. However, in many domains their use has been hampered by the practical difficulties of inference. This is particularly the case in computer vision, where models of the imaging process tend to be large, loopy and layered. For…

Cited by 5SourcePDFScholar