← Search

Peter Wonka

78 accepted papers

2026

Any Resolution Any Geometry: From Multi-View To Multi-Patch

CVPR 2026

Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high-resolution prediction remains difficult due to the trade-off between preserving fine local detail and maintaining global consistency. To address this challenge, we propose the Ultra Resolution Ge

Cited by 0SourcecodeScholar
2026

CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models

CVPR 2026

We introduce CADFS, a data-centric framework that enables large vision-language models to generate complex CAD design histories. Existing generative CAD systems are restricted to sketch-extrude operations due to simplified representations and limited datasets. We address this by introducing a Featur

Cited by 3SourcecodeScholar
2026

EasyV2V: A High-quality Instruction-based Video Editing Framework

CVPR 2026

While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization.We study the design space of data, architecture, and control, and introduce EasyV2V, a simple and effective framework for instruction-based video editing. On t

Cited by 4SourcecodeScholar
2026

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

CVPR 2026

Egocentric 3D human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method that addresses these challenges through two key contributi

Cited by 0SourceScholar
2026

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

ICML 2026poster

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large-language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes (e.g., kitchens, living rooms, bedrooms, bathrooms, and others), encoded symbolically in JSON or XML layouts. The …

Cited by 0SourceScholar
2026

LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

ICML 2026poster

We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional depth estimation, which is limited to visible surfaces, LaRI predicts multiple surfaces intersected by the camera rays using layered point maps. Compared to…

Cited by 0SourceScholar
2026

PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding

CVPR 2026

Current foundation models for 3D shapes excel at global tasks (retrieval, classification) but transfer poorly to local part-level reasoning. Recent approaches leverage vision and language foundation models to directly solve dense tasks through multi-view renderings and text queries. While promising,

Cited by 0SourcecodeScholar
2026

PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation

ICLR 2026poster

While current high-resolution depth estimation methods achieve strong results, they often suffer from computational inefficiencies due to reliance on heavyweight models and multiple inference steps, increasing inference time. To address this, we introduce PatchRefiner V2 (PRV2), which replaces heavy…

Cited by 0SourceScholar
2026

PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning

CVPR 2026

6D object pose estimation, which predicts the transformation of an object relative to the camera, remains challenging for unseen objects. Existing approaches typically rely on explicitly constructing feature correspondences between the query image and either the object model or template images. In t

Cited by 0SourcecodeScholar
2026

ShapeGen4D: Towards High Quality 4D Shape Generation from Videos

ICLR 2026poster

Video-conditioned 4D shape generation aims to recover time-varying 3D geometry and view-consistent appearance directly from an input video. In this work, we introduce a native video-to-4D shape generation framework that synthesizes a single dynamic 3D representation end-to-end from the video. Our…

Cited by 0SourceScholar
2025

4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion

CVPR 2025highlight

We propose 4Real-Video, a novel framework for generating 4D videos, organized as a grid of video frames with both time and viewpoint axes. In this grid, each row contains frames sharing the same timestep, while each column contains frames from the same viewpoint. One stream performs viewpoint updat…

Cited by 2SourcePDFScholar
2025

A3D: Does Diffusion Dream about 3D Alignment?

ICLR 2025poster

We tackle the problem of text-driven 3D generation from a geometry alignment perspective. Given a set of text prompts, we aim to generate a collection of objects with semantically corresponding parts aligned across them. Recent methods based on Score Distillation have succeeded in distilling the kno…

Cited by 0SourcePDFScholar
2025

Amodal Depth Anything: Amodal Depth Estimation in the Wild

ICCV 2025poster

Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive the geometry of occluded regions based on visible cues. Prior methods primarily rely on synthetic datasets and focus on m…

Cited by 0SourcePDFScholar
2025

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation

ICLR 2025poster

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding objects' placement and relationships from text descriptions. Nevertheless, existi…

2025

EditCLIP: Representation Learning for Image Editing

ICCV 2025poster

We introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve…

2025

Factored-NeuS: Reconstructing Surfaces, Illumination, and Materials of Possibly Glossy Objects

CVPR 2025poster

We develop a method that recovers the surface, materials, and illumination of a scene from its posed multi-view images. In contrast to prior work, it does not require any additional data and can handle glossy objects or bright lighting. It is a progressive inverse rendering approach, which consists…

Cited by 17SourcePDFScholar
2025

Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation

NeurIPS 2025poster

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze curren…

Cited by 0SourceScholar
2025

Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation

NeurIPS 2025spotlight

We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While diffusion model backbones are known to encode semantically rich…

Cited by 0SourceScholar
2025

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

ICCV 2025poster

We introduce the task of Language-Guided Object Placement in Real 3D Scenes. Given a 3D reconstructed point-cloud scene, a 3D asset, and a natural-language instruction, the goal is to place the asset so that the instruction is satisfied. The task demands tackling four intertwined challenges: (a) one…

Cited by 0SourcePDFScholar
2025

PrEditor3D: Fast and Precise 3D Shape Editing

CVPR 2025poster

We propose a training-free approach to 3D editing that enables the editing of a single shape and the reconstruction of a mesh within a few minutes. Leveraging 4-view images, user-guided text prompts, and rough 2D masks, our method produces an edited 3D mesh that aligns with the prompt. For this, our…

Cited by 3SourcePDFScholar
2025

T2Bs: Text-to-Character Blendshapes via Video Generation

ICCV 2025poster

We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion wi…

Cited by 0SourcePDFScholar
2025

V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video

ICCV 2025poster

We present V2M4, a novel 4D reconstruction method that directly generates a usable 4D mesh animation asset from a single monocular video. Unlike existing approaches that rely on priors from multi-view image and video generation models, our method is based on native 3D mesh generation models. Naively…

Cited by 0SourcePDFScholar
2025

VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models

CVPR 2025poster

We introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their deep understanding of image semantics. Yet, the majority…

Cited by 0SourcePDFScholar
2025

VoxelKP: A Voxel-based Network Architecture for Human Keypoint Estimation in LiDAR Data

ICCV 2025poster

We present VoxelKP, a novel fully sparse network architecture tailored for human keypoint estimation in LiDAR data. The key challenge is that objects are distributed sparsely in 3D space, while human keypoint detection requires detailed local information wherever humans are present. First, we introd…

2025

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

ICCV 2025poster

We propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP. Traditional methods for 3D keypoint detection rely heavily on ann…

Cited by 0SourcePDFScholar
2024

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

CVPR 2024poster

Recent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However current text-to-4D methods face a three-way tradeoff between the quality of scene appearance 3D structure and motion. For example text-to-image models and t…

2024

Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D Features

CVPR 2024poster

With the immense growth of dataset sizes and computing resources in recent years so-called foundation models have become popular in NLP and vision tasks. In this work we propose to explore foundation models for the task of keypoint detection on 3D shapes. A unique characteristic of keypoint detectio…

2024

Dissolving Is Amplifying: Towards Fine-Grained Anomaly Detection

ECCV 2024poster

"Medical imaging often contains critical fine-grained features, such as tumors or hemorrhages, which are crucial for diagnosis yet potentially too subtle for detection with conventional methods. In this paper, we introduce DIA, dissolving is amplifying. DIA is a fine-grained anomaly detection framew…

2024

Functional Diffusion

CVPR 2024poster

We propose functional diffusion a generative diffusion model focused on infinite-dimensional function data samples. In contrast to previous work functional diffusion works on samples that are represented by functions with a continuous domain. Functional diffusion can be seen as an extension of class…

Cited by 6SourcePDFScholar
2024

LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

ICLR 2024poster

Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these model…

2024

Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

ICLR 2024poster

We present ``Magic123'', a two-stage coarse-to-fine approach for high-quality, textured 3D mesh generation from a single image in the wild using *both 2D and 3D priors*. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-effic…

2024

PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth Estimation

CVPR 2024poster

Single image depth estimation is a foundational task in computer vision and generative modeling. However prevailing depth estimation models grapple with accommodating the increasing resolutions commonplace in today's consumer cameras and devices. Existing high-resolution strategies show promise but…

2024

PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation

ECCV 2024poster

"This paper introduces PatchRefiner, an advanced framework for metric single image depth estimation aimed at high-resolution real-domain inputs. While depth estimation is crucial for applications such as autonomous driving, 3D generative modeling, and 3D reconstruction, achieving accurate high-resol…

Cited by 6SourcePDFScholar
2024

Vivid-ZOO: Multi-View Video Generation with Diffusion Model

NeurIPS 2024poster

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeli…

Cited by 11SourcePDFScholar
2023

3D generation on ImageNet

ICLR 2023top-5%

All existing 3D-from-2D generators are designed for well-curated single-category datasets, where all the objects have (approximately) the same scale, 3D location, and orientation, and the camera always points to the center of the scene. This makes them inapplicable to diverse, in-the-wild datasets o…

2023

3DAvatarGAN: Bridging Domains for Personalized Editable Avatars

CVPR 2023poster

Modern 3D-GANs synthesize geometry and texture by training on large-scale datasets with a consistent structure. Training such models on stylized, artistic data, with often unknown, highly variable geometry, and camera information has not yet been shown possible. Can we train a 3D GAN on such artisti…

Cited by 48SourcePDFScholar
2023

SATR: Zero-Shot Semantic Segmentation of 3D Shapes

ICCV 2023poster

We explore the task of zero-shot semantic segmentation of 3D shapes by using large-scale off-the-shelf 2D im- age recognition models. Surprisingly, we find that modern zero-shot 2D object detectors are better suited for this task than contemporary text/image similarity predictors or even zero-shot 2…

Cited by 40PDFcodeScholar
2023

SLIBO-Net: Floorplan Reconstruction via Slicing Box Representation with Local Geometry Regularization

NeurIPS 2023poster

This paper focuses on improving the reconstruction of 2D floorplans from unstructured 3D point clouds. We identify opportunities for enhancement over the existing methods in three main areas: semantic quality, efficient representation, and local geometric details. To address these, we presents SLIBO…

Cited by 8SourcePDFScholar
2023

VIVE3D: Viewpoint-Independent Video Editing Using 3D-Aware GANs

CVPR 2023poster

We introduce VIVE3D, a novel approach that extends the capabilities of image-based 3D GANs to video editing and is able to represent the input video in an identity-preserving and temporally consistent way. We propose two new building blocks. First, we introduce a novel GAN inversion technique specif…

2022

3D CoMPaT: Composition of Materials on Parts of 3D Things

ECCV 2022poster

"We present 3D CoMPaT, a richly annotated large-scale dataset of more than 7.19 million rendered compositions of Materials on Parts of 7262 unique 3D Models; 990 compositions per model on average. 3D CoMPaT covers 43 shape categories, 235 unique part names, and 167 unique material classes that can b…

Cited by 16SourcePDFScholar
2022

HF-NeuS: Improved Surface Reconstruction Using High-Frequency Details

NeurIPS 2022accept

Neural rendering can be used to reconstruct implicit representations of shapes without 3D supervision. However, current neural surface reconstruction methods have difficulty learning high-frequency geometry details, so the reconstructed shapes are often over-smoothed. We develop HF-NeuS, a novel met…

2022

InsetGAN for Full-Body Image Generation

CVPR 2022poster

While GANs can produce photo-realistic images in ideal conditions for certain domains, the generation of full-body human images remains difficult due to the diversity of identities, hairstyles, clothing, and the variance in pose. Instead of modeling this complex domain with a single GAN, we propose…

Cited by 69PDFcodeScholar
2022

LocalBins: Improving Depth Estimation by Learning Local Distributions

ECCV 2022poster

"We propose a novel architecture for depth estimation from a single image. The architecture itself is based on the popular encoder-decoder architecture that is frequently used as a starting point for all dense regression tasks. We build on AdaBins which estimates a global distribution of depth value…

2022

Mind the Gap: Domain Gap Control for Single Shot Domain Adaptation for Generative Adversarial Networks

ICLR 2022poster

We present a new method for one shot domain adaptation. The input to our method is trained GAN that can produce images in domain A and a single reference image I_B from domain B. The proposed algorithm can translate any output of the trained GAN from domain A to domain B. There are two main advantag…

2022

On the Robustness of Quality Measures for GANs

ECCV 2022poster

"This work evaluates the robustness of quality measures of generative models such as Inception Score (IS) and Fréchet Inception Distance (FID). Analogous to the vulnerability of deep models against a variety of adversarial attacks, we show that such metrics can also be manipulated by additive pixel…

2021

Fast Sinkhorn Filters: Using Matrix Scaling for Non-Rigid Shape Correspondence With Functional Maps

CVPR 2021poster

In this paper, we provide a theoretical foundation for pointwise map recovery from functional maps and highlight its relation to a range of shape correspondence methods based on spectral alignment. With this analysis in hand, we develop a novel spectral registration technique: Fast Sinkhorn Filters,…

Cited by 64PDFcodeScholar
2021

Generative Layout Modeling Using Constraint Graphs

ICCV 2021poster

We propose a new generative model for layout generation. We generate layouts in three steps. First, we generate the layout elements as nodes in a layout graph. Second, we compute constraints between layout elements as edges in the layout graph. Third, we solve for the final layout using constrained…

Cited by 106PDFcodeScholar
2021

IntraTomo: Self-Supervised Learning-Based Tomography via Sinogram Synthesis and Prediction

ICCV 2021poster

We propose IntraTomo, a powerful framework that combines the benefits of learning-based and model-based approaches for solving highly ill-posed inverse problems in the Computed Tomography (CT) context. IntraTomo is composed of two core modules: a novel sinogram prediction module, and a geometry refi…

Cited by 106PDFcodeScholar
2021

SketchGen: Generating Constrained CAD Sketches

NeurIPS 2021poster

Computer-aided design (CAD) is the most widely used modeling approach for technical design. The typical starting point in these designs is 2D sketches which can later be extruded and combined to obtain complex three-dimensional assemblies. Such sketches are typically composed of parametric primitive…

Cited by 82SourcePDFScholar
2020

SEAN: Image Synthesis With Semantic Region-Adaptive Normalization

CVPR 2020oral

We propose semantic region-adaptive normalization (SEAN), a simple but effective building block for Generative Adversarial Networks conditioned on segmentation masks that describe the semantic regions in the desired output image. Using SEAN normalization, we can build a network architecture that can…

Cited by 731PDFcodeScholar
2020

TomoFluid: Reconstructing Dynamic Fluid From Sparse View Videos

CVPR 2020poster

Visible light tomography is a promising and increasingly popular technique for fluid imaging. However, the use of a sparse number of viewpoints in the capturing setups makes the reconstruction of fluid flows very challenging. In this paper, we present a state-of-the-art 4D tomographic reconstruction…

Cited by 37PDFScholar
2019

DuLa-Net: A Dual-Projection Network for Estimating Room Layouts From a Single RGB Panorama

CVPR 2019poster

We present a deep learning framework, called DuLa-Net, to predict Manhattan-world 3D room layouts from a single RGB panorama. To achieve better prediction accuracy, our method leverages two projections of the panorama at once, namely the equirectangular panorama-view and the perspective ceiling-vie…

Cited by 179PDFScholar
2018

Super-Resolution and Sparse View CT Reconstruction

ECCV 2018poster

We present a flexible framework for robust computed tomography (CT) reconstruction with a specific emphasis on recovering thin 1D and 2D manifolds embedded in 3D volumes. To reconstruct such structures at resolutions below the Nyquist limit of the CT image sensor, we devise a new 3D structure tensor…

Cited by 36SourcePDFScholar