← Search

Pascal Fua

93 accepted papers

2026

PhysGen: Physically Grounded 3D Shape Generation for Industrial Design

CVPR 2026

Existing generative models for 3D shapes can synthesize high-fidelity and visually plausible shapes. For certain classes of shapes that have undergone an engineering design process, the realism of the shape is tightly coupled with the underlying physical properties, e.g., aerodynamic efficiency for

Cited by 0SourcecodeScholar
2025

A View-consistent Sampling Method for Regularized Training of Neural Radiance Fields

ICCV 2025poster

Neural Radiance Fields (NeRF) has emerged as a compelling framework for scene representation and 3D recovery. To improve its performance on real-world data, depth regularizations have proven to be the most effective ones. However, depth estimation models not only require expensive 3D supervision in…

Cited by 0SourcePDFScholar
2025

IT$^3$: Idempotent Test-Time Training

ICML 2025poster

Deep learning models often struggle when deployed in real-world settings due to distribution shifts between training and test data. While existing approaches like domain adaptation and test-time training (TTT) offer partial solutions, they typically require additional data or domain-specific auxilia…

Cited by 0SourcePDFScholar
2025

LeFusion: Controllable Pathology Synthesis via Lesion-Focused Diffusion Models

ICLR 2025spotlight

Patient data from real-world clinical practice often suffers from data scarcity and long-tail imbalances, leading to biased outcomes or algorithmic unfairness. This study addresses these challenges by generating lesion-containing image-segmentation pairs from lesion-free images. Previous efforts in…

2025

Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures

CVPR 2025highlight

Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and…

Cited by 2SourcePDFScholar
2024

Enabling Uncertainty Estimation in Iterative Neural Networks

ICML 2024poster

Turning pass-through network architectures into iterative ones, which use their own output as input, is a well-known approach for boosting performance. In this paper, we argue that such architectures offer an additional benefit: The convergence rate of their successive outputs is highly correlated w…

2024

Garment Recovery with Shape and Deformation Priors

CVPR 2024poster

While modeling people wearing tight-fitting clothing has made great strides in recent years loose-fitting clothing remains a challenge. We propose a method that delivers realistic garment models from real-world images regardless of garment shape or deformation. To this end we introduce a fitting app…

2024

MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering

ECCV 2024poster

"Faithful human performance capture and free-view rendering from sparse RGB observations is a long-standing problem in Vision and Graphics. The main challenges are the lack of observations and the inherent ambiguities of the setting, e.g. occlusions and depth ambiguity. As a result, radiance fields,…

Cited by 9SourcePDFScholar
2024

Reconstruction of Manipulated Garment with Guided Deformation Prior

NeurIPS 2024poster

Modeling the shape of garments has received much attention, but most existing approaches assume the garments to be worn by someone, which constrains the range of shapes they can assume. In this work, we address shape recovery when garments are being manipulated instead of worn, which gives rise to a…

Cited by 1SourcePDFScholar
2023

DrapeNet: Garment Generation and Self-Supervised Draping

CVPR 2023poster

Recent approaches to drape garments quickly over arbitrary human bodies leverage self-supervision to eliminate the need for large training sets. However, they are designed to train one network per clothing item, which severely limits their generalization abilities. In our work, we rely on self-super…

2023

ISP: Multi-Layered Garment Draping with Implicit Sewing Patterns

NeurIPS 2023poster

Many approaches to draping individual garments on human body models are realistic, fast, and yield outputs that are differentiable with respect to the body shape on which they are draped. However, they are either unable to handle multi-layered clothing, which is prevalent in everyday dress, or restr…

2023

LightDepth: Single-View Depth Self-Supervision from Illumination Decline

ICCV 2023poster

Single-view depth estimation can be remarkably effective if there is enough ground-truth depth data for supervised training. However, there are scenarios, especially in medicine in the case of endoscopies, where such data cannot be obtained. In such cases, multi-view self-supervision and synthetic-t…

Cited by 8PDFScholar
2022

Adversarial Parametric Pose Prior

CVPR 2022oral

The Skinned Multi-Person Linear (SMPL) model represents human bodies by mapping pose and shape parameters to body meshes. However, not all pose and shape parameter values yield physically-plausible or even realistic body meshes. In other words, SMPL is under-constrained and may yield invalid results…

Cited by 44PDFcodeScholar
2022

ImplicitAtlas: Learning Deformable Shape Templates in Medical Imaging

CVPR 2022poster

Deep implicit shape models have become popular in the computer vision community at large but less so for biomedical applications. This is in part because large training databases do not exist and in part because biomedical annotations are often noisy. In this paper, we show that by introducing templ…

Cited by 35PDFScholar
2022

Learning To Align Sequential Actions in the Wild

CVPR 2022poster

State-of-the-art methods for self-supervised sequential action alignment rely on deep networks that find correspondences across videos in time. They either learn frame-to-frame mapping across sequences, which does not leverage temporal information, or assume monotonic alignment between each video pa…

Cited by 32PDFcodeScholar
2022

Learning to Simulate Realistic LiDARs

IROS 2022poster

Simulating realistic sensors is a challenging part in data generation for autonomous systems, often involving carefully handcrafted sensor design, scene properties, and physics modeling. To alleviate this, we introduce a pipeline for data-driven simulation of a realistic LiDAR sensor. We propose a m…

Cited by 19SourceScholar
2022

MeshUDF: Fast and Differentiable Meshing of Unsigned Distance Field Networks

ECCV 2022poster

"Unsigned Distance Fields (UDFs) can be used to represent non-watertight surfaces. However, current approaches to converting them into explicit meshes tend to either be expensive or to degrade the accuracy. Here, we extend the marching cube algorithm to handle UDFs, both fast and accurately. Moreove…

2022

Perspective Flow Aggregation for Data-Limited 6D Object Pose Estimation

ECCV 2022poster

"Most recent 6D object pose estimation methods, including unsupervised ones, require many real training images. Unfortunately, for some applications, such as those in space or deep under water, acquiring real images, even unannotated, is virtually impossible. In this paper, we propose a method that…

2021

Human Detection and Segmentation via Multi-View Consensus

ICCV 2021poster

Self-supervised detection and segmentation of foreground objects aims for accuracy without annotated training data. However, existing approaches predominantly rely on restrictive assumptions on appearance and motion. For scenes with dynamic activities and camera motion, we propose a multi-camera fra…

Cited by 3PDFcodeScholar
2021

PCLs: Geometry-Aware Neural Reconstruction of 3D Pose With Perspective Crop Layers

CVPR 2021poster

Local processing is an essential feature of CNNs and other neural network architectures -- it is one of the reasons why they work so well on images where relevant information is, to a large extent, local. However, perspective effects stemming from the projection in a conventional camera vary for dif…

Cited by 24PDFcodeScholar
2021

SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation

NeurIPS 2021poster

State-of-the-art semantic or instance segmentation deep neural networks (DNNs) are usually trained on a closed set of semantic classes. As such, they are ill-equipped to handle previously-unseen objects. However, detecting and localizing such objects is crucial for safety-critical applications such…

Cited by 155SourcecodeScholar
2021

Sketch2Mesh: Reconstructing and Editing 3D Shapes From Sketches

ICCV 2021poster

Reconstructing 3D shape from 2D sketches has long been an open problem because the sketches only provide very sparse and ambiguous information. In this paper, we use an encoder/decoder architecture for the sketch to mesh translation. When integrated into a user interface that provides camera paramet…

Cited by 87PDFScholar
2021

Temporally-Coherent Surface Reconstruction via Metric-Consistent Atlases

ICCV 2021poster

We propose a method for the unsupervised reconstruction of a temporally-coherent sequence of surfaces from a sequence of time-evolving point clouds, yielding dense, semantically meaningful correspondences between all keyframes. We represent the reconstructed surface as an atlas, using a neural netwo…

Cited by 7PDFScholar
2021

Wide-Depth-Range 6D Object Pose Estimation in Space

CVPR 2021poster

6D pose estimation in space poses unique challenges that are not commonly encountered in the terrestrial setting. One of the most striking differences is the lack of atmospheric scattering, allowing objects to be visible from a great distance while complicating illumination conditions. Currently ava…

Cited by 104PDFScholar
2020

ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion Capture

CVPR 2020oral

The accuracy of monocular 3D human pose estimation depends on the viewpoint from which the image is captured. While freely moving cameras, such as on drones, provide control over this viewpoint, automatically positioning them at the location which will yield the highest accuracy remains an open prob…

Cited by 47PDFcodeScholar
2020

DISK: Learning local features with policy gradient

NeurIPS 2020spotlight

Local feature frameworks are difficult to learn in an end-to-end fashion due to the discreteness inherent to the selection and matching of sparse keypoints. We introduce DISK (DIScrete Keypoints), a novel method that overcomes these obstacles by leveraging principles from Reinforcement Learning (RL)…

2020

Deformation-Aware Unpaired Image Translation for Pose Estimation on Laboratory Animals

CVPR 2020poster

Our goal is to capture the pose of real animals using synthetic training examples, without using any manual supervision. Our focus is on neuroscience model organisms, to be able to study how neural circuits orchestrate behaviour. Human pose estimation attains remarkable accuracy when trained on real…

Cited by 51PDFScholar
2020

Lightweight Multi-View 3D Pose Estimation Through Camera-Disentangled Representation

CVPR 2020poster

We present a lightweight solution to recover 3D pose from multi-view images captured with spatially calibrated cameras. Building upon recent advances in interpretable representation learning, we exploit 3D geometry to fuse input images into a unified latent representation of pose, which is disentang…

Cited by 147PDFScholar
2020

MeshSDF: Differentiable Iso-Surface Extraction

NeurIPS 2020spotlight

Geometric Deep Learning has recently made striking progress with the advent of continuous Deep Implicit Fields. They allow for detailed modeling of watertight surfaces of arbitrary topology while not relying on a 3D Euclidean grid, resulting in a learnable parameterization that is not limited in res…

2020

Shape Reconstruction by Learning Differentiable Surface Representations

CVPR 2020poster

Generative models that produce point clouds have emerged as a powerful tool to represent 3D surfaces, and the best current ones rely on learning an ensemble of parametric representations. Unfortunately, they offer no control over the deformations of the surface patches that form the ensemble and thu…

Cited by 66PDFcodeScholar
2020

Towards Reliable Evaluation of Algorithms for Road Network Reconstruction from Aerial Images

ECCV 2020poster

Existing connectivity-oriented performance measures rank road delineation algorithms inconsistently, which makes it difficult to decide which one is best for a given application. We show that these inconsistencies stem from design flaws that make the metrics insensitive to whole classes of errors. T…

Cited by 10SourcePDFScholar
2019

Backpropagation-Friendly Eigendecomposition

NeurIPS 2019poster

Eigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it with the Power Iteration method, particularly when dealing with large matrices. While this can be mitigated by partitio…

2019

Beyond Cartesian Representations for Local Descriptors

ICCV 2019poster

The dominant approach for learning local patch descriptors relies on small image regions whose scale must be properly estimated a priori by a keypoint detector. In other words, if two patches are not in correspondence, their descriptors will not match. A strategy often used to alleviate this problem…

Cited by 135PDFcodeScholar
2019

GarNet: A Two-Stream Network for Fast and Accurate 3D Cloth Draping

ICCV 2019poster

While Physics-Based Simulation (PBS) can accurately drape a 3D garment on a 3D body, it remains too costly for real-time applications, such as virtual try-on. By contrast, inference in a deep network, requiring a single forward pass, is much faster. Taking advantage of this, we propose a novel archi…

Cited by 154PDFScholar
2019

Geometric and Physical Constraints for Drone-Based Head Plane Crowd Density Estimation

IROS 2019poster

State-of-the-art methods for counting people in crowded scenes rely on deep networks to estimate crowd density in the image plane. While useful for this purpose, this image-plane density has no immediate physical meaning because it is subject to perspective distortion. This is a concern in sequences…

Cited by 63SourceScholar
2019

Neural Scene Decomposition for Multi-Person Motion Capture

CVPR 2019poster

Learning general image representations has proven key to the success of many computer vision tasks. For example, many approaches to image understanding problems rely on deep networks that were initially trained on ImageNet, mostly because the learned features are a valuable starting point to learn f…

Cited by 61PDFScholar
2019

Recurrent U-Net for Resource-Constrained Segmentation

ICCV 2019poster

State-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on standard GPUs. In this paper, we introduce a novel recurrent U-Net architecture that preserves the compactness of the origi…

Cited by 135PDFScholar
2018

Beyond the Pixel-Wise Loss for Topology-Aware Delineation

CVPR 2018poster

Delineation of curvilinear structures is an important problem in Computer Vision with multiple practical applications. With the advent of Deep Learning, many current approaches on automatic delineation have focused on finding more powerful deep architectures, but have continued using the habitual pi…

Cited by 301SourcePDFScholar
2018

Eigendecomposition-free Training of Deep Networks with Zero Eigenvalue-based Losses

ECCV 2018poster

Many classical Computer Vision problems, such as essential matrix computation and pose estimation from 3D to 2D correspondences, can be solved by finding the eigenvector corresponding to the smallest, or zero, eigenvalue of a matrix representing a linear system. Incorporating this in deep learning f…

Cited by 54SourcePDFScholar
2018

Every Smile Is Unique: Landmark-Guided Diverse Smile Generation

CVPR 2018poster

Each smile is unique: one person surely smiles in different ways (e.g., closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this one-to-many video generation problem, we propose a novel deep le…

Cited by 82SourcePDFScholar
2018

FishEyeRecNet: A Multi-Context Collaborative Deep Network for Fisheye Image Rectification

ECCV 2018poster

Images captured by sheye lenses violate the pinhole camera assumption and suer from distortions. Rectication of sheye images is therefore a crucial preprocessing step for many computer vision applications. In this paper, we propose an end-to-end multi-context collaborative deep network for removing…

Cited by 163SourcePDFScholar
2018

Learning Monocular 3D Human Pose Estimation From Multi-View Images

CVPR 2018poster

Accurate 3D human pose estimation from single images is possible with sophisticated deep-net architectures that have been trained on very large datasets. However, this still leaves open the problem of capturing motions for which no such database exists. Manual annotation is tedious, slow, and error…

Cited by 308SourcePDFScholar
2018

Learning to Find Good Correspondences

CVPR 2018poster

We develop a deep architecture to learn to find good correspondences for wide-baseline stereo. Given a set of putative sparse matches and the camera intrinsics, we train our network in an end-to-end fashion to label the correspondences as inliers or outliers, while simultaneously using them to recov…

Cited by 685SourcePDFScholar
2018

Modeling Facial Geometry Using Compositional VAEs

CVPR 2018poster

We propose a method for learning non-linear face geometry representations using deep generative models. Our model is a variational autoencoder with multiple levels of hidden variables where lower layers capture global geometry and higher ones encode more local deformations. Based on that, we pr…

Cited by 151SourcePDFScholar
2018

Unsupervised Geometry-Aware Representation for 3D Human Pose Estimation

ECCV 2018poster

Modern 3D human pose estimation techniques rely on deep networks, which require large amounts of training data. While weakly-supervised methods require less supervision, by utilizing 2D poses or multi-view imagery without annotations, they still need a sufficiently large set of samples with 3D annot…

Cited by 313SourcePDFScholar
2018

WILDTRACK: A Multi-Camera HD Dataset for Dense Unscripted Pedestrian Detection

CVPR 2018poster

People detection methods are highly sensitive to occlusions between pedestrians, which are extremely frequent in many situations where cameras have to be mounted at a limited height. The reduction of camera prices allows for the generalization of static multi-camera set-ups. Using joint visual infor…

2017

Flight Dynamics-Based Recovery of a UAV Trajectory Using Ground Cameras

CVPR 2017oral

We propose a new method to estimate the 6-dof trajectory of a flying object such as a quadrotor UAV within a 3D airspace monitored using multiple fixed ground cameras. It is based on a new structure from motion formulation for the 3D reconstruction of a single moving point with known motion dynamics…

Cited by 45PDFcodeScholar
2017

Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation

ICCV 2017poster

Most recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D joint locations from which 3D coordinates are inferred. Both approaches have their strengths and weaknesses and we therefo…

Cited by 327PDFScholar
2017

Social Scene Understanding: End-To-End Multi-Person Action Localization and Collective Activity Recognition

CVPR 2017oral

We present a unified framework for understanding human social behaviors in raw image sequences. Our model jointly detects multiple individuals, infers their social actions, and estimates the collective actions with a single feed-forward pass through a neural network. We propose a single architecture…

Cited by 296PDFScholar
2016

Active Learning for Delineation of Curvilinear Structures

CVPR 2016poster

Many recent delineation techniques owe much of their increased effectiveness to path classification algorithms that make it possible to distinguish promising paths from others. The downside of this development is that they require annotated training data, which is tedious to produce. In this…

Cited by 19PDFScholar
2016

Direct Prediction of 3D Body Poses From Motion Compensated Sequences

CVPR 2016poster

We propose an efficient approach to exploiting motion information from consecutive frames of a video sequence to recover the 3D pose of people. Previous approaches typically compute candidate poses in individual frames and then link them in a post-processing step to resolve ambiguities. By contrast…

Cited by 273PDFScholar
2016

Learning to Match Aerial Images With Deep Attentive Architectures

CVPR 2016poster

Image matching is a fundamental problem in Computer Vision. In the context of feature-based matching, SIFT and its variants have long excelled in a wide array of applications. However, for ultra-wide baselines, as in the case of aerial images captured under large camera rotations, the appearance var…

Cited by 93PDFScholar
2016

Vision-based Unmanned Aerial Vehicle detection and tracking for sense and avoid systems

IROS 2016poster

We propose an approach for on-line detection of small Unmanned Aerial Vehicles (UAVs) and estimation of their relative positions and velocities in the 3D environment from a single moving camera in the context of sense and avoid systems. This problem is challenging both from a detection point of view…

Cited by 76SourceScholar
2015

A Novel Representation of Parts for Accurate 3D Object Detection and Tracking in Monocular Images

ICCV 2015poster

We present a method that estimates in real-time and under challenging conditions the 3D pose of a known object. Our method relies only on grayscale images since depth cameras fail on metallic objects; it can handle poorly textured objects, and cluttered, changing environments; the pose it predi…

Cited by 135PDFScholar
2015

Dense Image Registration and Deformable Surface Reconstruction in Presence of Occlusions and Minimal Texture

ICCV 2015poster

Deformable surface tracking from monocular images is well-known to be under-constrained. Occlusions often make the task even more challenging, and can result in failure if the surface is not sufficiently textured. In this work, we explicitly address the problem of 3D reconstruction of poorly texture…

Cited by 53PDFScholar
2015

Discriminative Learning of Deep Convolutional Feature Point Descriptors

ICCV 2015poster

Deep learning has revolutionalized image-level tasks such as classification, but patch-level tasks, such as correspondence, still rely on hand-crafted features, e.g. SIFT. In this paper we use Convolutional Neural Networks (CNNs) to learn discriminant patch representations and in particular train a…

Cited by 1022PDFcodeScholar
2015

Hot or Not: Exploring Correlations Between Appearance and Temperature

ICCV 2015poster

In this paper we explore interactions between the appearance of an outdoor scene and the ambient temperature. By studying statistical correlations between image sequences from outdoor cameras and temperature measurements we identify two interesting interactions. First, semantically meaningful region…

Cited by 32PDFScholar
2015

Kullback-Leibler Proximal Variational Inference

NeurIPS 2015poster

We propose a new variational inference method based on the Kullback-Leibler (KL) proximal term. We make two contributions towards improving efficiency of variational inference. Firstly, we derive a KL proximal-point algorithm and show its equivalence to gradient descent with natural gradient in stoc…

Cited by 55SourcePDFScholar