← Search

Philippe Weinzaepfel

36 accepted papers

2025

DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers

CVPR 2025poster

Recent multi-teacher distillation methods have unified the encoders of multiple foundation models into a single encoder, achieving competitive performance on core vision tasks like classification, segmentation, and depth estimation. This led us to ask: Could similar success be achieved when the pool…

Cited by 0SourcePDFScholar
2025

HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction

ICCV 2025poster

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly predicting dense scene geometry, they are primarily trained…

Cited by 0SourcePDFScholar
2025

Kinaema: a recurrent sequence model for memory and pose in motion

NeurIPS 2025poster

One key aspect of spatially aware robots is the ability to "find their bearings", ie. to correctly situate themselves or previously seen spaces. In this work, we focus on this particular scenario of continuous robotics operations, where information observed before an actual episode start is exploite…

Cited by 0SourceScholar
2025

Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors

CVPR 2025poster

We present Pow3R, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors at test time, Pow3R incorporates any combination of auxiliary information such a…

Cited by 2SourcePDFScholar
2024

"PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation"

ECCV 2024poster

"Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of human-centric vision, albeit CLIP-like representations encode…

Cited by 1SourcePDFScholar
2024

Cross-view and Cross-pose Completion for 3D Human Understanding

CVPR 2024poster

Human perception and understanding is a major domain of computer vision which like many other vision subdomains recently stands to gain from the use of large models pre-trained on large datasets. We hypothesize that the most common pre-training strategy of relying on general purpose object-centric i…

Cited by 5SourcePDFScholar
2024

End-to-End (Instance)-Image Goal Navigation through Correspondence as an Emergent Phenomenon

ICLR 2024poster

Most recent work in goal oriented visual navigation resorts to large-scale machine learning in simulated environments. The main challenge lies in learning compact representations generalizable to unseen environments and in learning high-capacity perception modules capable of reasoning on high-dimens…

Cited by 9SourcePDFScholar
2024

Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot

ECCV 2024poster

"We present , a strong model for multi-person 3D human mesh recovery from a single RGB image. Predictions encompass the whole body, , including hands and facial expressions, using the SMPL-X parametric model and 3D location in the camera coordinate system. Our model detects people by predicting coar…

2024

UNIC: Universal Classification Models via Multi-teacher Distillation

ECCV 2024poster

"Pretrained models have become a commodity and offer strong results on a broad range of tasks. In this work, we focus on classification and seek to learn a unique encoder able to take from several complementary pretrained models. We aim at even stronger generalization across a variety of classificat…

Cited by 6SourcePDFScholar
2024

Weatherproofing Retrieval for Localization with Generative AI and Geometric Consistency

ICLR 2024poster

State-of-the-art visual localization approaches generally rely on a first image retrieval step whose role is crucial. Yet, retrieval often struggles when facing varying conditions, due to e.g. weather or time of day, with dramatic consequences on the visual localization accuracy. In this paper, we i…

Cited by 0SourcePDFScholar
2024

Win-Win: Training High-Resolution Vision Transformers from Two Windows

ICLR 2024poster

Transformers have become the standard in state-of-the-art vision architectures, achieving impressive performance on both image-level and dense pixelwise tasks. However, training vision transformers for high-resolution pixelwise tasks has a prohibitive cost. Typical solutions boil down to hierarchica…

Cited by 4SourcePDFScholar
2023

CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

ICCV 2023poster

Despite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow. The application of self-supervised concepts, such as instance discrimination or masked image mode…

Cited by 100PDFcodeScholar
2023

PoseFix: Correcting 3D Human Poses with Natural Language

ICCV 2023poster

Automatically producing instructions to modify one's posture could open the door to endless applications, such as personalized coaching and in-home physical therapy. Tackling the reverse problem (i.e., refining a 3D pose based on some natural language feedback) could help for assisted 3D character a…

Cited by 29PDFScholar
2022

Barely-Supervised Learning: Semi-supervised Learning with Very Few Labeled Images

AAAI 2022technical

This paper tackles the problem of semi-supervised learning when the set of labeled samples is limited to a small number of images per class, typically less than 10, problem that we refer to as barely-supervised learning. We analyze in depth the behavior of a state-of-the-art semi-supervised method,…

Cited by 32SourcePDFScholar
2022

CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

NeurIPS 2022accept

Masked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible patches as sole input. This pre-training leads to state-of-the-…

2022

Learning Super-Features for Image Retrieval

ICLR 2022poster

Methods that combine local and global features have recently shown excellent performance on multiple challenging deep image retrieval benchmarks, but their use of local features raises at least two issues. First, these local features simply boil down to the localized map activations of a neural netw…

2022

PUMP: Pyramidal and Uniqueness Matching Priors for Unsupervised Learning of Local Descriptors

CVPR 2022poster

Existing approaches for learning local image descriptors have shown remarkable achievements in a wide range of geometric tasks. However, most of them require per-pixel correspondence-level supervision, which is difficult to acquire at scale and in high quality. In this paper, we propose to explicitl…

Cited by 16PDFcodeScholar
2022

PoseGPT: Quantization-Based 3D Human Motion Generation and Forecasting

ECCV 2022poster

"We address the problem of action-conditioned generation of human motion sequences. Existing work falls into two categories: forecast models conditioned on observed past motions, or generative models conditioned action labels and duration only. In contrast, we generate motion conditioned on observat…

2022

PoseScript: 3D Human Poses from Natural Language

ECCV 2022poster

"Natural language is leveraged in many computer vision tasks such as image captioning, cross-modal retrieval or visual question answering, to provide fine-grained semantic information. While human pose is key to human understanding, current 3D human pose datasets lack detailed language descriptions.…

Cited by 66SourcePDFScholar
2021

Large-Scale Localization Datasets in Crowded Indoor Spaces

CVPR 2021poster

Estimating the precise location of a camera using visual localization enables interesting applications such as augmented reality or robot navigation. This is particularly useful in indoor environments where other localization technologies, such as GNSS, fail. Indoor spaces impose interesting challen…

Cited by 50PDFcodeScholar
2021

Multi-FinGAN: Generative Coarse-To-Fine Sampling of Multi-Finger Grasps

ICRA 2021poster

While there exists many methods for manipulating rigid objects with parallel-jaw grippers, grasping with multi-finger robotic hands remains a quite unexplored research topic. Reasoning and planning collision-free trajectories on the additional degrees of freedom of several fingers represents an impo…

Cited by 63SourcecodeScholar
2020

DOPE: Distillation Of Part Experts for whole-body 3D pose estimation in the wild

ECCV 2020poster

We introduce DOPE, the first method to detect and estimate whole-body 3D human poses, including bodies, hands and faces, in the wild. Achieving this level of details is key for a number of applications that require understanding the interactions of the people with each other or with the environment.…

Cited by 65SourcePDFScholar
2020

Hard Negative Mixing for Contrastive Learning

NeurIPS 2020poster

Contrastive learning has become a key component of self-supervised learning approaches for computer vision. By learning to embed two augmented versions of the same image close to each other and to push the embeddings of different images apart, one can train highly transferable visual representations…

2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

SuperLoss: A Generic Loss for Robust Curriculum Learning

NeurIPS 2020poster

Curriculum learning is a technique to improve a model performance and generalization based on the idea that easy samples should be presented before difficult ones during training. While it is generally complex to estimate a priori the difficulty of a given sample, recent works have shown that curric…

2019

MARS: Motion-Augmented RGB Stream for Action Recognition

CVPR 2019poster

Most state-of-the-art methods for action recognition consist of a two-stream architecture with 3D convolutions: an appearance stream for RGB frames and a motion stream for optical flow frames. Although combining flow with RGB improves the performance, the cost of computing accurate optical flow is…

Cited by 336PDFScholar
2019

R2D2: Reliable and Repeatable Detector and Descriptor

NeurIPS 2019oral

Interest point detection and local feature description are fundamental steps in many computer vision applications. Classical approaches are based on a detect-then-describe paradigm where separate handcrafted methods are used to first identify repeatable keypoints and then represent them with a local…

2019

Visual Localization by Learning Objects-Of-Interest Dense Match Regression

CVPR 2019poster

We introduce a novel CNN-based approach for visual localization from a single RGB image that relies on densely matching a set of Objects-of-Interest (OOIs). In this paper, we focus on planar objects which are highly descriptive in an environment, such as paintings in museums or logos and storefronts…

Cited by 53PDFScholar
2018

PoTion: Pose MoTion Representation for Action Recognition

CVPR 2018poster

Most state-of-the-art methods for action recognition rely on a two-stream architecture that processes appearance and motion independently. In this paper, we claim that considering them jointly offers rich information for action recognition. We introduce a novel representation that gracefully encodes…

Cited by 381SourcePDFScholar
2017

Action Tubelet Detector for Spatio-Temporal Action Localization

ICCV 2017poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level that are then linked or tracked across time. In this paper, we leverage the temporal continuity of videos instead of operating at the frame level. We propose the ACtion Tubelet detector…

Cited by 437PDFcodeScholar
2015

EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow

CVPR 2015poster

We propose a novel approach for optical flow estimation, targeted at large displacements with significant occlusions. It consists of two steps: i) dense matching by edge-preserving interpolation from a sparse set of matches; ii) variational energy minimization initialized with the dense matches. The…

2015

Learning to Detect Motion Boundaries

CVPR 2015poster

We propose a learning-based approach for motion boundary detection. Precise localization of motion boundaries is essential for the success of optical flow estimation, as motion boundaries correspond to discontinuities of the optical flow field. The proposed approach allows to predict motion boundari…