← Search

Leonidas Guibas

112 accepted papers

2026

DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-To-Sim Physical Calibration

ICRA 2026poster

Simulating optical tactile sensors presents significant challenges due to their high deformability and intricate optical properties. To address these issues and enable a physically accurate simulation, we propose DOT-Sim: Differentiable Optical Tactile Simulation. Unlike prior simulators that rely o…

2026

Dynamic Reflections: Probing Video Representations with Text Alignment

ICLR 2026poster

The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal natur…

Cited by 0SourceScholar
2026

FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

CVPR 2026

We present FunREC, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunREC operates directly on in-t

Cited by 0SourcecodeScholar
2026

GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation Policies

ICML 2026poster

Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action abstractions require understanding three-dimensional geometric transformations. Here, we introduce GeoMoLa (Geometry-Aware Motion Latents), which learns discrete …

Cited by 0SourceScholar
2026

Mixture of Contexts for Long Video Generation

ICLR 2026poster

Long video generation is fundamentally a long context memory problem: models must retain and retrieve salient events across a long range without collapsing or drifting. However, scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-atte…

Cited by 0SourceScholar
2026

Mode Seeking meets Mean Seeking for Long Video Generation

ICML 2026poster

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to narrow domains. While multi-resolution image training works because higher resolution is largely an interpolation of th…

Cited by 7SourceScholar
2026

RINO: Rotation-Invariant Non-Rigid Correspondences

CVPR 2026

Dense 3D shape correspondence remains a central challenge in computer vision and graphics as many deep learning approaches still rely on intermediate geometric features or handcrafted descriptors, limiting their effectiveness under non-isometric deformations, partial data, and non-manifold inputs. T

Cited by 0SourceScholar
2026

Rodrigues Network for Learning Robot Actions

ICLR 2026oral

Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers lack inductive biases that reflect the underlying kinematic structure of articulated systems. To this end, we propose the **Neural Rodrigues Operator**, a lear…

Cited by 0SourceScholar
2026

SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling

ICLR 2026poster

Generative methods for 3D assets have recently achieved remarkable progress, yet providing intuitive and precise control over the object geometry remains a key challenge. Existing approaches predominantly rely on text or image prompts, which often fall short in geometric specificity: language can be…

Cited by 0SourceScholar
2026

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

ICML 2026poster

Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled data from external vision sources or synthetic…

Cited by 0SourceScholar
2026

pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

ICLR 2026poster

Few-step diffusion or flow-based generative models typically distill a velocity-predicting teacher into a student that predicts a shortcut towards denoised data. This format mismatch has led to complex distillation procedures that often suffer from a quality--diversity trade-off. To address this, we…

Cited by 0SourcecodeScholar
2025

ARCH: Hierarchical Hybrid Learning for Long-Horizon Contact-Rich Robotic Assembly

CoRL 2025poster

Generalizable long-horizon robotic assembly requires reasoning at multiple levels of abstraction. While end-to-end imitation learning (IL) is a promising approach, it typically requires large amounts of expert demonstration data and often struggles to achieve the high precision demanded by assembly…

Cited by 0SourceScholar
2025

AllTracker: Efficient Dense Point Tracking at High Resolution

ICCV 2025poster

We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing point tracking methods, our approach delivers high-resolution and dense (all-pixel) correspondence fields, which can be…

2025

BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing

CVPR 2025highlight

3D graphics editing is crucial in applications like movie production and game design, yet it remains a time-consuming process that demands highly specialized domain expertise. Automating this process is challenging because graphical editing requires performing a variety of tasks, each requiring dist…

2025

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

CVPR 2025poster

Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts, i.e., "identity-preserving generation". This setting, along with many other ta…

Cited by 11SourcePDFScholar
2025

Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

CVPR 2025poster

Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic operations with complex 3D/4D scenes remains challenging. This dif…

Cited by 1SourcePDFScholar
2025

FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement

CVPR 2025highlight

Scene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Models (MLLMs) excel at semantic tasks, their application to 3D scene generation is hindered by their limited grounding on…

Cited by 2SourcePDFScholar
2025

Gaussian Mixture Flow Matching Models

ICML 2025poster

Diffusion models approximate the denoising distribution as a Gaussian and predict its mean, whereas flow matching models reparameterize the Gaussian mean as flow velocity. However, they underperform in few-step sampling due to discretization error and tend to produce over-saturated colors under clas…

2025

Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large Motion

ICCV 2025poster

Existing dynamic scene interpolation methods typically assume that the motion between consecutive timesteps is small enough so that displacements can be locally approximated by linear models. In practice, even slight deviations from this small-motion assumption can cause conventional techniques to f…

Cited by 0SourcePDFScholar
2025

GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling

CVPR 2025poster

We present GroomLight, a novel method for relightable hair appearance modeling from multi-view images. Existing hair capture methods struggle to balance photorealistic rendering with relighting capabilities. Analytical material models, while physically grounded, often fail to fully capture appearanc…

2025

HouseLayout3D: A Benchmark and Training-free Baseline for 3D Layout Estimation in the Wild

NeurIPS 2025poster

Current 3D layout estimation models are predominantly trained on synthetic datasets biased toward simplistic, single-floor scenes. This prevents them from generalizing to complex, multi-floor buildings, often forcing a per-floor processing approach that sacrifices global context. Few works have atte…

Cited by 0SourceScholar
2025

InfoGS: Efficient Structure-Aware 3D Gaussians via Lightweight Information Shaping

ICLR 2025poster

3D Gaussians, as an explicit scene representation, typically involve thousands to millions of elements per scene. This makes it challenging to control the scene in ways that reflect the underlying semantics, where the number of independent entities is typically much smaller. Especially, if one wants…

2025

MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps

ICCV 2025poster

This paper addresses the challenge of learning semantically and functionally meaningful 3D motion priors from real-world videos, in order to enable prediction of future 3D scene motion from a single input image. We propose a novel pixel-aligned Motion Map (MoMap) representation for 3D scene motion,…

Cited by 0SourcePDFScholar
2025

MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

CVPR 2025highlight

We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundati…

2025

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

ICLR 2025poster

Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D…

2025

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

ICCV 2025poster

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for env…

Cited by 0SourcePDFScholar
2025

Robot Learning from Any Images

CoRL 2025poster

We introduce RoLA, a framework that transforms any in‑the‑wild image into an interactive, physics‑enabled robotic environment. Unlike previous methods, RoLA operates directly on a single image without requiring additional hardware or digital assets. Our framework democratizes robotic data generatio…

Cited by 0SourcecodeScholar
2025

SceneCrafter: Controllable Multi-View Driving Scene Editing

CVPR 2025poster

Simulation is crucial for developing and evaluating autonomous vehicle (AV) systems. Recent literature builds on a new generation of generative models to synthesize highly realistic images for full-stack simulation. However, purely synthetically generated scenes are not grounded in reality and have…

Cited by 0SourcePDFScholar
2025

Self-Calibrating Gaussian Splatting for Large Field-of-View Reconstruction

ICCV 2025poster

Large field-of-view (FOV) cameras can simplify and accelerate scene capture because they provide complete coverage with fewer views. However, existing reconstruction pipelines fail to take full advantage of large-FOV input data because they convert input views to perspective images, resulting in str…

2025

SplatTalk: 3D VQA with Gaussian Splatting

ICCV 2025poster

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D vision-language models (VLMs) have achieved remarkable success in 2D V…

Cited by 0SourcePDFScholar
2025

SuperDec: 3D Scene Decomposition with Superquadrics Primitives

ICCV 2025poster

We present SuperDec, an approach for compact 3D scene representations based on geometric primitives, namely superquadrics.While most recent works leverage geometric primitives to obtain photorealistic 3D scene representations, we propose to leverage them to obtain a compact yet expressive representa…

Cited by 0SourcePDFScholar
2025

Video Perception Models for 3D Scene Synthesis

NeurIPS 2025poster

Automating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or…

Cited by 0SourceScholar
2025

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

ICCV 2025poster

We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in temporal changes. Specifically, we aim to capture frequent co-occurring changes ("trends") across a city over a certain p…

Cited by 0SourcePDFScholar
2024

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

CVPR 2024poster

Recent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However current text-to-4D methods face a three-way tradeoff between the quality of scene appearance 3D structure and motion. For example text-to-image models and t…

2024

ActAnywhere: Subject-Aware Video Background Generation

NeurIPS 2024poster

We study a novel problem to automatically generate video background that tailors to foreground subject motion. It is an important problem for the movie industry and visual effects community, which traditionally requires tedious manual efforts to solve. To this end, we propose ActAnywhere, a video di…

2024

ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling

ECCV 2024poster

"Recognizing and disentangling visual attributes from objects is a foundation to many computer vision applications. While large vision-language representations like CLIP had largely resolved the task of zero-shot object recognition, zero-shot visual attribute recognition remains a challenge because…

2024

BlenderAlchemy: Editing 3D Graphics with Vision-Language Models

ECCV 2024poster

"Graphics design is important for various applications, including movie production and game design. To create a high-quality scene, designers usually need to spend hours in software like Blender, in which they might need to interleave and repeat operations, such as connecting material nodes, hundred…

2024

CAD: Photorealistic 3D Generation via Adversarial Distillation

CVPR 2024poster

The increased demand for 3D data in AR/VR robotics and gaming applications gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation Sampling (SDS) algorithm to optimize a 3D representation such that the rendered i…

Cited by 13SourcePDFScholar
2024

Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control

NeurIPS 2024poster

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent approaches that condition video generation models on camera traj…

Cited by 24SourcePDFScholar
2024

ConDense: Consistent 2D-3D Pre-training for Dense and Sparse Features from Multi-View Images

ECCV 2024oral

"To advance the state of the art in the creation of 3D foundation models, this paper introduces the framework for 3D pre-training utilizing existing pre-trained 2D networks and large-scale multi-view datasets. We propose a novel 2D-3D joint training scheme to extract co-embedded 2D and 3D features i…

Cited by 5SourcePDFScholar
2024

D$^3$RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation

CoRL 2024poster

Depth sensing is an important problem for 3D vision-based robotics. Yet, a real-world active stereo or ToF depth camera often produces noisy and incomplete depth which bottlenecks robot performances. In this work, we propose D3RoMa, a learning-based depth estimation framework on stereo image pairs t…

Cited by 4SourceScholar
2024

Denoising Vision Transformers

ECCV 2024oral

"We study a crucial yet often overlooked issue inherent to Vision Transformers (ViTs): feature maps of these models exhibit grid-like artifacts (“Original features” in fig:teaser), which hurt the performance of ViTs in downstream dense prediction tasks such as semantic segmentation, depth prediction…

2024

EquivAct: SIM(3)-Equivariant Visuomotor Policies beyond Rigid Object Manipulation

ICRA 2024poster

If a robot masters folding a kitchen towel, we would expect it to master folding a large beach towel. However, existing policy learning methods that rely on data augmentation still don’t guarantee such generalization. Our insight is to add equivariance to both the visual object representation and po…

Cited by 38SourcecodeScholar
2024

GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation

CVPR 2024poster

Despite recent advances in text-to-3D generative methods there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each such as how well the asset aligned with the input text. These metrics lack the flexibility to generalize to different evaluati…

2024

MultiPhys: Multi-Person Physics-aware 3D Motion Estimation

CVPR 2024poster

We introduce MultiPhys a method designed for recovering multi-person motion from monocular videos. Our focus lies in capturing coherent spatial placement between pairs of individuals across varying degrees of engagement. MultiPhys being physically aware exhibits robustness to jittering and occlusion…

Cited by 5SourcePDFScholar
2024

NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis

CVPR 2024poster

We address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a specific object which outputs the distance to the valid interaction manifold given a human pose as input. This interaction field…

Cited by 44SourcePDFScholar
2024

Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping

CoRL 2024poster

One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences across 3D scenes, their features are point-based and restricted to object surfaces, l…

Cited by 2SourceScholar
2024

PACE: Pose Annotations in Cluttered Environments

ECCV 2024poster

"We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered scenarios. PACE provides a large-scale real-world benchmark for both instance-level and category-level settings. The ben…

2024

PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations

ECCV 2024poster

"[width=0.9]figure/teaserv 4.pdf Figure 1: PhysAvatar is a novel framework that captures the physics of dressed 3D avatars from visual observations, enabling a wide spectrum of applications, such as (a) animation, (b) relighting, and (c) redressing, with high-fidelity rendering results."

2024

Probing the 3D Awareness of Visual Foundation Models

CVPR 2024poster

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task their intermediate representations are useful for other visual tasks such as detection and segmentation. Given…

2024

ProvNeRF: Modeling per Point Provenance in NeRFs as a Stochastic Field

NeurIPS 2024poster

Neural radiance fields (NeRFs) have gained popularity with multiple works showing promising results across various applications. However, to the best of our knowledge, existing works do not explicitly model the distribution of training camera poses, or consequently the triangulation quality, a key f…

2024

RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation

CoRL 2024poster

This work proposes a retrieve-and-transfer framework for zero-shot robotic manipulation, dubbed RAM, featuring generalizability across various objects, environments, and embodiments. Unlike existing approaches that learn manipulation from expensive in-domain demonstrations, RAM capitalizes on a retr…

Cited by 26SourcecodeScholar
2024

SAGE: Bridging Semantic and Actionable Parts for GEneralizable Articulated-Object Manipulation under Language Instructions

RSS 2024poster

To interact with daily-life articulated objects of diverse structures and functionalities, understanding the object parts plays a central role in both user instruction comprehension and task execution. However, the possible discordance between the semantic meaning and physics functionalities of the…

Cited by 0SourcePDFScholar
2024

SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation

ICLR 2024poster

Humans demonstrate remarkable skill in transferring manipulation abilities across objects of varying shapes, poses, and appearances, a capability rooted in their understanding of semantic correspondences between different instances. To equip robots with a similar high-level comprehension, we present…

Cited by 16SourcePDFScholar
2024

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

CVPR 2024poster

Understanding and reasoning about spatial relationships is crucial for Visual Question Answering (VQA) and robotics. Vision Language Models (VLMs) have shown impressive performance in some VQA benchmarks but struggle with 3D spatial reasoning such as recognizing distances or size differences between…

Cited by 198SourcePDFScholar
2024

Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models

ICRA 2024poster

Object tracking is central to robot perception and scene understanding, allowing robots to parse a video stream in terms of moving objects with names. Tracking-by-detection has long been a dominant paradigm for object tracking of specific object categories [1], [2]. Recently, large-scale pre-trained…

Cited by 12SourceScholar
2023

ALTO: Alternating Latent Topologies for Implicit 3D Reconstruction

CVPR 2023poster

This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatial arrangement of latent encodings is important to recover detail. One school of thought is to encode a latent vector for…

Cited by 35SourcePDFScholar
2023

Affection: Learning Affective Explanations for Real-World Visual Data

CVPR 2023poster

In this work, we explore the space of emotional reactions induced by real-world images. For this, we first introduce a large-scale dataset that contains both categorical emotional reactions and free-form textual explanations for 85,007 publicly available images, analyzed by 6,283 annotators who were…

Cited by 18SourcePDFScholar
2023

Banana: Banach Fixed-Point Network for Pointcloud Segmentation with Inter-Part Equivariance

NeurIPS 2023spotlight

Equivariance has gained strong interest as a desirable network property that inherently ensures robust generalization. However, when dealing with complex systems such as articulated objects or multi-object scenes, effectively capturing inter-part transformations poses a challenge, as it becomes enta…

Cited by 15SourcePDFScholar
2023

CC3D: Layout-Conditioned Generation of Compositional 3D Scenes

ICCV 2023poster

In this work, we introduce CC3D, a conditional generative model that synthesizes complex 3D scenes conditioned on 2D semantic scene layouts, trained using single-view images. Different from most existing 3D GANs that limit their applicability to aligned single objects, we focus on generating complex…

Cited by 44PDFScholar
2023

DiffFacto: Controllable Part-Based 3D Point Cloud Generation with Cross Diffusion

ICCV 2023poster

While the community of 3D point cloud generation has witnessed a big growth in recent years, there still lacks an effective way to enable intuitive user control in the generation process, hence limiting the general utility of such methods. Since an intuitive way of decomposing a shape is through its…

Cited by 29PDFScholar
2023

EFEM: Equivariant Neural Field Expectation Maximization for 3D Object Segmentation Without Scene Supervision

CVPR 2023poster

We introduce Equivariant Neural Field Expectation Maximization (EFEM), a simple, effective, and robust geometric algorithm that can segment objects in 3D scenes without annotations or training on scenes. We achieve such unsupervised segmentation by exploiting single object shape priors. We make two…

Cited by 22SourcePDFScholar
2023

FindThis: Language-Driven Object Disambiguation in Indoor Environments

CoRL 2023poster

Natural language is naturally ambiguous. In this work, we consider interactions between a user and a mobile service robot tasked with locating a desired object, specified by a language utterance. We present a task FindThis, which addresses the problem of how to disambiguate and locate the particular…

Cited by 11SourceScholar
2023

GINA-3D: Learning To Generate Implicit Neural Assets in the Wild

CVPR 2023poster

Modeling the 3D world from sensor data for simulation is a scalable way of developing testing and validation environments for robotic learning problems such as autonomous driving. However, manually creating or re-creating real-world-like environments is difficult, expensive, and not scalable. Recent…

Cited by 21SourcePDFScholar
2023

Generating Part-Aware Editable 3D Shapes Without 3D Supervision

CVPR 2023poster

Impressive progress in generative models and implicit representations gave rise to methods that can generate 3D shapes of high quality. However, being able to locally control and edit shapes is another essential property that can unlock several content creation applications. Local control can be ach…

2023

JacobiNeRF: NeRF Shaping With Mutual Information Gradients

CVPR 2023poster

We propose a method that trains a neural radiance field (NeRF) to encode not only the appearance of the scene but also semantic correlations between scene points, regions, or entities -- aiming to capture their mutual co-variation patterns. In contrast to the traditional first-order photometric reco…

2023

LEGO-Net: Learning Regular Rearrangements of Objects in Rooms

CVPR 2023poster

Humans universally dislike the task of cleaning up a messy room. If machines were to help us with this task, they must understand human criteria for regular arrangements, such as several types of symmetry, co-linearity or co-circularity, spacing uniformity in linear or circular patterns, and further…

Cited by 62SourcePDFScholar
2023

MetaCLUE: Towards Comprehensive Visual Metaphors Research

CVPR 2023poster

Creativity is an indispensable part of human cognition and also an inherent part of how we make sense of the world. Metaphorical abstraction is fundamental in communicating creative ideas through nuanced relationships between abstract concepts such as feelings. While computer vision benchmarks and a…

2023

NAP: Neural 3D Articulated Object Prior

NeurIPS 2023poster

We propose Neural 3D Articulated object Prior (NAP), the first 3D deep generative model to synthesize 3D articulated object models. Despite the extensive research on generating 3D static objects, compositions, or scenes, there are hardly any approaches on capturing the distribution of articulated ob…

Cited by 16SourcePDFScholar
2023

NeRDi: Single-View NeRF Synthesis With Language-Guided Diffusion As General Image Priors

CVPR 2023poster

2D-to-3D reconstruction is an ill-posed problem, yet humans are good at solving this problem due to their prior knowledge of the 3D world developed over years. Driven by this observation, we propose NeRDi, a single-view NeRF synthesis framework with general image priors from 2D diffusion models. For…

Cited by 169SourcePDFScholar
2023

NeRF Revisited: Fixing Quadrature Instability in Volume Rendering

NeurIPS 2023poster

Neural radiance fields (NeRF) rely on volume rendering to synthesize novel views. Volume rendering requires evaluating an integral along each ray, which is numerically approximated with a finite sum that corresponds to the exact integral along the ray under piecewise constant volume density. As a co…

2023

Nerflets: Local Radiance Fields for Efficient Structure-Aware 3D Scene Representation From 2D Supervision

CVPR 2023poster

We address efficient and structure-aware 3D scene representation from images. Nerflets are our key contribution-- a set of local neural radiance fields that together represent a scene. Each nerflet maintains its own spatial position, orientation, and extent, within which it contributes to panoptic,…

Cited by 54SourcePDFScholar
2023

SCADE: NeRFs from Space Carving With Ambiguity-Aware Depth Estimates

CVPR 2023poster

Neural radiance fields (NeRFs) have enabled high fidelity 3D reconstruction from multiple 2D input views. However, a well-known drawback of NeRFs is the less-than-ideal performance under a small number of views, due to insufficient constraints enforced by volumetric rendering. To address this issue,…

2023

ShapeTalk: A Language Dataset and Framework for 3D Shape Edits and Deformations

CVPR 2023poster

Editing 3D geometry is a challenging task requiring specialized skills. In this work, we aim to facilitate the task of editing the geometry of 3D models through the use of natural language. For example, we may want to modify a 3D chair model to "make its legs thinner" or to "open a hole in its back"…

2023

SinGRAF: Learning a 3D Generative Radiance Field for a Single Scene

CVPR 2023poster

Generative models have shown great promise in synthesizing photorealistic 3D objects, but they require large amounts of training data. We introduce SinGRAF, a 3D-aware generative model that is trained with a few input images of a single scene. Once trained, SinGRAF generates different realizations o…

Cited by 16SourcePDFScholar
2023

Towards Learning Geometric Eigen-Lengths Crucial for Fitting Tasks

ICML 2023poster

Some extremely low-dimensional yet crucial geometric eigen-lengths often determine the success of some geometric tasks. For example, the *height* of an object is important to measure to check if it can fit between the shelves of a cabinet, while the *width* of a couch is crucial when trying to move…

Cited by 5SourcePDFScholar
2023

VDN-NeRF: Resolving Shape-Radiance Ambiguity via View-Dependence Normalization

CVPR 2023poster

We propose VDN-NeRF, a method to train neural radiance fields (NeRFs) for better geometry under non-Lambertian surface and dynamic lighting conditions that cause significant variation in the radiance of a point when viewed from different angles. Instead of explicitly modeling the underlying factors…

2022

IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor Scenes

ICLR 2022poster

Building embodied intelligent agents that can interact with 3D indoor environments has received increasing research attention in recent years. While most works focus on single-object or agent-object visual functionality and affordances, our work proposes to study a novel, underexplored, kind of visu…

Cited by 7SourcePDFScholar
2022

LADIS: Language Disentanglement for 3D Shape Editing

EMNLP 2022finding

Natural language interaction is a promising direction for democratizing 3D shape design. However, existing methods for text-driven 3D shape editing face challenges in producing decoupled, local edits to 3D shapes. We address this problem by learning disentangled latent representations that ground la…

2022

NeuForm: Adaptive Overfitting for Neural Shape Editing

NeurIPS 2022accept

Neural representations are popular for representing shapes as they can be used for data cleanup, model completion, shape editing, and shape synthesis. Current neural representations can be categorized as either overfitting to a single object instance, or representing a collection of objects. However…

Cited by 17SourcePDFScholar
2022

Object Pursuit: Building a Space of Objects via Discriminative Weight Generation

ICLR 2022poster

We propose a framework to continuously learn object-centric representations for visual learning and understanding. Existing object-centric representations either rely on supervisions that individualize objects in the scene, or perform unsupervised disentanglement that can hardly deal with complex sc…

2022

Object Scene Representation Transformer

NeurIPS 2022accept

A compositional understanding of the world in terms of objects and their geometry in 3D space is considered a cornerstone of human cognition. Facilitating the learning of such a representation in neural networks holds promise for substantially improving labeled data efficiency. As a key step in this…

Cited by 116SourcePDFScholar
2022

VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects

ICLR 2022poster

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their myriad semantic categories, diverse shape geometry, and complicat…

Cited by 104SourcePDFScholar
2021

Intrinsic Dimension, Persistent Homology and Generalization in Neural Networks

NeurIPS 2021poster

Disobeying the classical wisdom of statistical learning theory, modern deep neural networks generalize well even though they typically contain millions of parameters. Recently, it has been shown that the trajectories of iterative optimization algorithms can possess \emph{fractal structures}, and the…

2021

Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point Clouds

NeurIPS 2021poster

Category-level object pose estimation aims to find 6D object poses of previously unseen object instances from known categories without access to object CAD models. To reduce the huge amount of pose annotations needed for category-level learning, we propose for the first time a self-supervised learni…

2021

O2O-Afford: Annotation-Free Large-Scale Object-Object Affordance Learning

CoRL 2021poster

Contrary to the vast literature in modeling, perceiving, and understanding agent-object (e.g., human-object, hand-object, robot-object) interaction in computer vision and robotics, very few past works have studied the task of object-object interaction, which also plays an important role in robotic m…

Cited by 73SourceScholar
2021

SketchGen: Generating Constrained CAD Sketches

NeurIPS 2021poster

Computer-aided design (CAD) is the most widely used modeling approach for technical design. The typical starting point in these designs is 2D sketches which can later be extruded and combined to obtain complex three-dimensional assemblies. Such sketches are typically composed of parametric primitive…

Cited by 82SourcePDFScholar
2020

6D Camera Relocalization in Ambiguous Scenes via Continuous Multimodal Inference

ECCV 2020poster

We present a multimodal camera relocalization framework that captures ambiguities and uncertainties with continuous mixture models defined on the manifold of camera poses. In highly ambiguous environments, which can easily arise due to symmetries and repetitive structures in the scene, computing one…

2020

CaSPR: Learning Canonical Spatiotemporal Point Cloud Representations

NeurIPS 2020spotlight

We propose CaSPR, a method to learn object-centric Canonical Spatiotemporal Point Cloud Representations of dynamically moving or evolving objects. Our goal is to enable information aggregation over time and the interrogation of object state at any spatiotemporal neighborhood in the past, observed or…

2020

Deformation-Aware 3D Model Embedding and Retrieval

ECCV 2020poster

We introduce a new problem of mph{retrieving} 3D models that are mph{deformable} to a given query shape and present a novel deep mph{deformation-aware} embedding to solve this retrieval task. 3D model retrieval is a fundamental operation for recovering a clean and complete 3D model from a noisy and…

2020

Generative 3D Part Assembly via Dynamic Graph Learning

NeurIPS 2020poster

Autonomous part assembly is a challenging yet crucial task in 3D computer vision and robotics. Analogous to buying an IKEA furniture, given a set of 3D parts that can assemble a single shape, an intelligent agent needs to perceive the 3D part geometry, reason to propose pose estimations for the inpu…

Cited by 100SourcePDFScholar
2020

Learning 3D Part Assembly from a Single Image

ECCV 2020poster

Autonomous assembly is a crucial capability for robots in many applications. For this task, several problems such as obstacle avoidance, motion planning, and actuator control have been extensively studied in robotics. However, when it comes to task specification, the space of possibilities remains u…

2020

PT2PC: Learning to Generate 3D Point Cloud Shapes from Part Tree Conditions

ECCV 2020poster

Generative 3D shape modeling is a fundamental research area in computer vision and interactive computer graphics, with many real-world applications. This paper investigates the novel problem of generating a 3D point cloud geometry for a shape from a symbolic part tree representation. In order to lea…

Cited by 49SourcePDFScholar
2020

PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding

ECCV 2020poster

Arguably one of the top success stories of deep learning is transfer learning. The finding that pre-training a network on a rich source set (g, ImageNet) can help boost performance once fine-tuned on a usually much smaller target set, has been instrumental to many applications in language and vision…

2020

Quaternion Equivariant Capsule Networks for 3D Point Clouds

ECCV 2020poster

We present a 3D capsule module for processing point clouds that is equivariant to 3D rotations and translations, as well as invariant to permutations of the input points. The operator receives a sparse set of local reference frames, computed from an input point cloud and establishes end-to-end trans…

Cited by 112SourcePDFScholar
2020

ReferIt3D: Neural Listeners for Fine-Grained 3D Object Identification in Real-World Scenes

ECCV 2020poster

In this work we study the problem of using referential language to identify common objects in real-world 3D scenes. We focus on a challenging setup where the referred object belongs to a extit{fine-grained} object class and the underlying scene contains extit{multiple} object instances of that class…

2020

ShapeFlow: Learnable Deformation Flows Among 3D Shapes

NeurIPS 2020spotlight

We present ShapeFlow, a flow-based model for learning a deformation space for entire classes of 3D shapes with large intra-class variations. ShapeFlow allows learning a multi-template deformation space that is agnostic to shape topology, yet preserves fine geometric details. Different from a generat…

Cited by 101SourcePDFScholar
2020

Side-Tuning: A Baseline for Network Adaptation via Additive Side Networks

ECCV 2020poster

When training a neural network for a desired task, one may prefer to adapt a pre-trained network rather than starting from randomly initialized weights. Adaptation can be useful in cases when training data is scarce, when a single learner needs to perform multiple tasks, or when one wishes to encode…

Cited by 257SourcePDFScholar
2020

Towards Precise Completion of Deformable Shapes

ECCV 2020poster

According to Aristotle, a philosopher in Ancient Greece, {\it ``the whole is greater than the sum of its parts''}. This statement was adopted to explain human perception by the Gestalt psychology school of thought in the twentieth century. Here, we claim that observing part of an object which was pr…

2020

Which Tasks Should Be Learned Together in Multi-task Learning?

ICML 2020poster

Many computer vision applications require solving multiple tasks in real-time. A neural network can be trained to solve multiple tasks simultaneously using multi-task learning. This can save computation at inference time as only a single network needs to be evaluated. Unfortunately, this often leads…

2019

A Condition Number for Joint Optimization of Cycle-Consistent Networks

NeurIPS 2019spotlight

A recent trend in optimizing maps such as dense correspondences between objects or neural networks between pairs of domains is to optimize them jointly. In this context, there is a natural \textsl{cycle-consistency} constraint, which regularizes composite maps associated with cycles, i.e., they are…

2019

Learning to Navigate Using Mid-Level Visual Priors

CoRL 2019

How much does having visual priors about the world (e.g. the fact that the world is 3D) assist in learning to perform downstream motor tasks (e.g. navigating a complex environment)? What are the consequences of not utilizing such visual priors in learning? We study these questions by integrating a g

2019

Multiview Aggregation for Learning Category-Specific Shape Reconstruction

NeurIPS 2019poster

We investigate the problem of learning category-specific 3D shape reconstruction from a variable number of RGB views of previously unobserved object instances. Most approaches for multiview shape reconstruction operate on sparse shape representations, or assume a fixed number of views. We present a…

2019

PeerNets: Exploiting Peer Wisdom Against Adversarial Attacks

ICLR 2019poster

Deep learning systems have become ubiquitous in many aspects of our lives. Unfortunately, it has been shown that such systems are vulnerable to adversarial attacks, making them prone to potential unlawful uses. Designing deep neural networks that are robust to adversarial attacks is a fundamental s…

Cited by 0SourcePDFScholar
2018

Deep Functional Dictionaries: Learning Consistent Semantic Structures on 3D Models from Functions

NeurIPS 2018poster

Various 3D semantic attributes such as segmentation masks, geometric features, keypoints, and materials can be encoded as per-point probe functions on 3D geometries. Given a collection of related 3D shapes, we consider how to jointly analyze such probe functions over different shapes, and how to dis…

2018

Learning Representations and Generative Models for 3D Point Clouds

ICLR 2018workshop

Three-dimensional geometric data offer an excellent domain for studying representation learning and generative modeling. In this paper, we look at geometric data represented as point clouds. We introduce a deep autoencoder (AE) network with excellent reconstruction quality and generalization ability…

Cited by 1765SourceScholar
2018

Learning Representations and Generative Models for 3D Point Clouds

ICML 2018oral

Three-dimensional geometric data offer an excellent domain for studying representation learning and generative modeling. In this paper, we look at geometric data represented as point clouds. We introduce a deep AutoEncoder (AE) network with state-of-the-art reconstruction quality and generalization…

2017

PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space

NeurIPS 2017poster

Few prior works study deep learning on point sets. PointNet is a pioneer in this direction. However, by design PointNet does not capture local structures induced by the metric space points live in, limiting its ability to recognize fine-grained patterns and generalizability to complex scenes. In thi…

Cited by 14360SourcePDFScholar
2016

FPNN: Field Probing Neural Networks for 3D Data

NeurIPS 2016poster

Building discriminative representations for 3D data has been an important task in computer graphics and computer vision research. Convolutional Neural Networks (CNNs) have shown to operate on 2D images with great success for a variety of tasks. Lifting convolution operators to 3D (3DCNNs) seems like…

2015

Deep Knowledge Tracing

NeurIPS 2015poster

Knowledge tracing, where a machine models the knowledge of a student as they interact with coursework, is an established and significantly unsolved problem in computer supported education.In this paper we explore the benefit of using recurrent neural networks to model student learning.This family of…

2015

Learning Program Embeddings to Propagate Feedback on Student Code

ICML 2015poster

Providing feedback, both assessing final work and giving hints to stuck students, is difficult for open-ended assignments in massive online classes which can range from thousands to millions of students. We introduce a neural network method to encode programs as a linear mapping from an embedded pre…

Cited by 249SourcePDFScholar