← Search

Diane Larlus

27 accepted papers

2026

Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

CVPR 2026

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key contribution is a compact latent space that encodes the entire anim

Cited by 0SourcecodeScholar
2025

DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers

CVPR 2025poster

Recent multi-teacher distillation methods have unified the encoders of multiple foundation models into a single encoder, achieving competitive performance on core vision tasks like classification, segmentation, and depth estimation. This led us to ask: Could similar success be achieved when the pool…

Cited by 0SourcePDFScholar
2025

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

ICCV 2025poster

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only synthetic data while generalizing well to real data in a zero-…

2025

LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting Scenes

ICCV 2025poster

We address the problem of extending the capabilities of vision foundation models such as DINO, SAM, and CLIP, to 3D tasks. Specifically, we introduce a novel method to uplift 2D image features into Gaussian Splatting representations of 3D scenes. Unlike traditional approaches that rely on minimizing…

Cited by 0SourcePDFScholar
2025

Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos

CVPR 2025poster

Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown that 3D techniques can at last improve outputs extracted from…

Cited by 0SourcePDFScholar
2024

UNIC: Universal Classification Models via Multi-teacher Distillation

ECCV 2024poster

"Pretrained models have become a commodity and offer strong results on a broad range of tasks. In this work, we focus on classification and seek to learn a unique encoder able to take from several complementary pretrained models. We aim at even stronger generalization across a variety of classificat…

Cited by 6SourcePDFScholar
2024

Weatherproofing Retrieval for Localization with Generative AI and Geometric Consistency

ICLR 2024poster

State-of-the-art visual localization approaches generally rely on a first image retrieval step whose role is crucial. Yet, retrieval often struggles when facing varying conditions, due to e.g. weather or time of day, with dramatic consequences on the visual localization accuracy. In this paper, we i…

Cited by 0SourcePDFScholar
2023

EPIC Fields: Marrying 3D Geometry and Video Understanding

NeurIPS 2023poster

Neural rendering is fuelling a unification of learning, 3D geometry and video understanding that has been waiting for more than two decades. Progress, however, is still hampered by a lack of suitable datasets and benchmarks. To address this gap, we introduce EPIC Fields, an augmentation of EPIC-KITC…

2023

Fake It Till You Make It: Learning Transferable Representations From Synthetic ImageNet Clones

CVPR 2023poster

Recent image generation models such as Stable Diffusion have exhibited an impressive ability to generate fairly realistic images starting from a simple text prompt. Could such models render real images obsolete for training image prediction models? In this paper, we answer part of this provocative q…

Cited by 179SourcePDFScholar
2023

No Reason for No Supervision: Improved Generalization in Supervised Models

ICLR 2023top-25%

We consider the problem of training a deep neural network on a given classification task, e.g., ImageNet-1K (IN1K), so that it excels at both the training task as well as at other (future) transfer tasks. These two seemingly contradictory properties impose a trade-off between improving the model’s g…

Cited by 35SourcePDFScholar
2023

SLACK: Stable Learning of Augmentations With Cold-Start and KL Regularization

CVPR 2023poster

Data augmentation is known to improve the generalization capabilities of neural networks, provided that the set of transformations is chosen with care, a selection often performed manually. Automatic data augmentation aims at automating this process. However, most recent approaches still rely on som…

Cited by 6SourcePDFScholar
2022

ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity

ICLR 2022poster

An intuitive way to search for images is to use queries composed of an example image and a complementary text. While the first provides rich and implicit context for the search, the latter explicitly calls for new traits, or specifies how some elements of the example image should be changed to retri…

2022

Granularity-Aware Adaptation for Image Retrieval over Multiple Tasks

ECCV 2022poster

"Strong image search models can be learned for a specific domain, ie. set of labels, provided that some labeled images of that domain are available. A practical visual search model, however, should be versatile enough to solve multiple retrieval tasks simultaneously, even if those cover very differe…

Cited by 9SourcePDFScholar
2022

Learning Super-Features for Image Retrieval

ICLR 2022poster

Methods that combine local and global features have recently shown excellent performance on multiple challenging deep image retrieval benchmarks, but their use of local features raises at least two issues. First, these local features simply boil down to the localized map activations of a neural netw…

2022

On the Road to Online Adaptation for Semantic Image Segmentation

CVPR 2022poster

We propose a new problem formulation and a corresponding evaluation framework to advance research on unsupervised domain adaptation for semantic image segmentation. The overall goal is fostering the development of adaptive learning systems that will continuously learn, without supervision, in ever-c…

Cited by 35PDFcodeScholar
2021

Concept Generalization in Visual Representation Learning

ICCV 2021poster

Measuring concept generalization, i.e., the extent to which models trained on a set of (seen) visual concepts can be leveraged to recognize a new set of (unseen) concepts, is a popular way of evaluating visual representations, especially in a self-supervised learning framework. Nonetheless, the choi…

Cited by 50PDFcodeScholar
2021

Continual Adaptation of Visual Representations via Domain Randomization and Meta-Learning

CVPR 2021poster

Most standard learning approaches lead to fragile models which are prone to drift when sequentially trained on samples of a different nature -- the well-known "catastrophic forgetting" issue. In particular, when a model consecutively learns from different visual domains, it tends to forget the past…

Cited by 96PDFcodeScholar
2021

Probabilistic Embeddings for Cross-Modal Retrieval

CVPR 2021poster

Cross-modal retrieval methods build a common representation space for samples from multiple modalities, typically from the vision and the language domains. For images and their captions, the multiplicity of the correspondences makes the task particularly challenging. Given an image (respectively a c…

Cited by 275PDFcodeScholar
2020

Hard Negative Mixing for Contrastive Learning

NeurIPS 2020poster

Contrastive learning has become a key component of self-supervised learning approaches for computer vision. By learning to embed two augmented versions of the same image close to each other and to push the embeddings of different images apart, one can train highly transferable visual representations…

2019

Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings

ICCV 2019poster

We address the problem of cross-modal fine-grained action retrieval between text and video. Cross-modal retrieval is commonly achieved through learning a shared embedding space, that can indifferently embed modalities. In this paper, we propose to enrich the embedding by disentangling parts-of-speec…

Cited by 184PDFScholar
2018

Self-Supervised Learning of Geometrically Stable Features Through Probabilistic Introspection

CVPR 2018poster

Self-supervision can dramatically cut back the amount of manually-labelled data required to train deep neural networks. While self-supervision has usually been considered for tasks such as image classification, in this paper we aim at extending it to geometry-oriented tasks such as semantic matching…

Cited by 87SourcePDFScholar
2018

Semi-convolutional Operators for Instance Segmentation

ECCV 2018poster

Object detection and instance segmentation are dominated by region-based methods such as Mask RCNN. However, there is a growing interest in reducing these problems to pixel labeling tasks, as the latter could be more efficient, could be integrated seamlessly in image-to-image network architectures a…

Cited by 110SourcePDFScholar
2017

AnchorNet: A Weakly Supervised Network to Learn Geometry-Sensitive Features for Semantic Matching

CVPR 2017poster

Despite significant progress of deep learning in recent years, state-of-the-art semantic matching methods still rely on legacy features such as SIFT or HoG. We argue that the strong invariance properties that are key to the success of recent deep architectures on the classification task make them un…

Cited by 71PDFScholar
2017

Beyond Instance-Level Image Retrieval: Leveraging Captions to Learn a Global Visual Representation for Semantic Retrieval

CVPR 2017poster

Querying with an example image is a simple and intuitive interface to retrieve information from a visual database. Most of the research in image retrieval has focused on the task of instance-level image retrieval, where the goal is to retrieve images that contain the same object instance as the quer…

Cited by 113PDFScholar