← Search

Patrick Pérez

38 accepted papers

2026

Vision-Speech Models: Teaching Speech Models to Converse about Images

CVPR 2026

The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards building a multimodal speech model able to freely converse about images. Building such a conversational Vision-Speech mod

Cited by 0SourcecodeScholar
2024

A Simple Recipe for Language-guided Domain Generalized Segmentation

CVPR 2024poster

Generalization to new domains not seen during training is one of the long-standing challenges in deploying neural networks in real-world applications. Existing generalization techniques either necessitate external images for augmentation and/or aim at learning invariant representations by imposing v…

2024

CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

ECCV 2024poster

"The popular CLIP model displays impressive zero-shot capabilities thanks to its seamless interaction with arbitrary text prompts. However, its lack of spatial awareness makes it unsuitable for dense computer vision tasks, e.g., semantic segmentation, without an additional fine-tuning step that ofte…

2024

PointBeV: A Sparse Approach for BeV Predictions

CVPR 2024poster

Bird's-eye View (BeV) representations have emerged as the de-facto shared space in driving applications offering a unified space for sensor data fusion and supporting various downstream tasks. However conventional models use grids with fixed resolution and range and face computational inefficiencies…

2024

Reliability in Semantic Segmentation: Can We Use Synthetic Data?

ECCV 2024poster

"Assessing the robustness of perception models to covariate shifts and their ability to detect out-of-distribution (OOD) inputs is crucial for safety-critical applications such as autonomous vehicles. By nature of such applications, however, the relevant data is difficult to collect and annotate. In…

2024

Three Pillars Improving Vision Foundation Model Distillation for Lidar

CVPR 2024poster

Self-supervised image backbones can be used to address complex 2D tasks (e.g. semantic segmentation object discovery) very efficiently and with little or no downstream supervision. Ideally 3D backbones for lidar should be able to inherit these properties after distillation of these powerful 2D featu…

2024

Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?

ICRA 2024poster

Motion forecasting is crucial in enabling autonomous vehicles to anticipate the future trajectories of surrounding agents. To do so, it requires solving mapping, detection, tracking, and then forecasting problems, in a multi-step pipeline. In this complex system, advances in conventional forecasting…

Cited by 19SourcecodeScholar
2023

OCTET: Object-Aware Counterfactual Explanations

CVPR 2023poster

Nowadays, deep vision models are being widely deployed in safety-critical applications, e.g., autonomous driving, and explainability of such models is becoming a pressing concern. Among explanation methods, counterfactual explanations aim to find minimal and interpretable changes to the input image…

2023

PODA: Prompt-driven Zero-shot Domain Adaptation

ICCV 2023poster

Domain adaptation has been vastly investigated in computer vision but still requires access to target images at train time, which might be intractable in some uncommon conditions. In this paper, we propose the task of 'Prompt-driven Zero-shot Domain Adaptation', where we adapt a model trained on a s…

Cited by 62PDFcodeScholar
2023

T-UDA: Temporal Unsupervised Domain Adaptation in Sequential Point Clouds

IROS 2023poster

Deep perception models have to reliably cope with an open-world setting of domain shifts induced by different geographic regions, sensor properties, mounting positions, and several other reasons. Since covering all domains with annotated data is technically intractable due to the endless possible va…

Cited by 4SourcecodeScholar
2023

Teachers in Concordance for Pseudo-Labeling of 3D Sequential Data

RA-L 2023

Automatic pseudo-labeling is a powerful tool to tap into large amounts of sequential unlabeled data. It is especially appealing in safety-critical applications of autonomous driving, where performance requirements are extreme, datasets are large, and manual labeling is very challenging. We propose t

Cited by 7SourcecodeScholar
2023

Unsupervised Object Localization: Observing the Background To Discover Objects

CVPR 2023poster

Recent advances in self-supervised visual representation learning have paved the way for unsupervised methods tackling tasks such as object discovery and instance segmentation. However, discovering objects in an image with no supervision is a very hard task; what are the desired objects, when to sep…

2022

Active Learning Strategies for Weakly-Supervised Object Detection

ECCV 2022poster

"Object detectors trained with weak annotations are affordable alternatives to fully-supervised counterparts. However, there is still a significant performance gap between them. We propose to narrow this gap by fine-tuning a base pre-trained weakly-supervised detector with a few fully-annotated samp…

2022

Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-Modal Distillation

ECCV 2022poster

"This work investigates learning pixel-wise semantic image segmentation in urban scenes without any manual annotation, just from the raw non-curated data collected by cars which, equipped with cameras and LiDAR sensors, drive around a city. Our contributions are threefold. First, we propose a novel…

2022

HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance

ECCV 2022poster

"Marker-less monocular 3D human motion capture (MoCap) with scene interactions is a challenging research topic relevant for extended reality, robotics and virtual avatar generation. Due to the inherent depth ambiguity of monocular settings, 3D motions captured with existing methods often contain sev…

Cited by 31SourcePDFScholar
2022

STEEX: Steering Counterfactual Explanations with Semantics

ECCV 2022poster

"As deep learning models are increasingly used in safety-critical applications, explainability and trustworthiness become major concerns. For simple images, such as low-resolution face portraits, synthesizing visual counterfactual explanations has recently been proposed as a way to uncover the decis…

2021

Artificial Dummies for Urban Dataset Augmentation

AAAI 2021technical

Existing datasets for training pedestrian detectors in images suffer from limited appearance and pose variation. The most challenging scenarios are rarely included because they are too difficult to capture due to safety reasons, or they are very unlikely to happen. The strict safety requirements in…

2021

Multi-Target Adversarial Frameworks for Domain Adaptation in Semantic Segmentation

ICCV 2021poster

In this work, we address the task of unsupervised domain adaptation (UDA) for semantic segmentation in presence of multiple target domains: the objective is to train a single model that can handle all these domains at test time. Such a multi-target adaptation is crucial for a variety of scenarios th…

Cited by 42PDFcodeScholar
2021

Multi-View Radar Semantic Segmentation

ICCV 2021poster

Understanding the scene around the ego-vehicle is key to assisted and autonomous driving. Nowadays, this is mostly conducted using cameras and laser scanners, despite their reduced performances in adverse weather conditions. Automotive radars are low-cost active sensors that measure properties of su…

Cited by 95PDFcodeScholar
2020

QuEST: Quantized Embedding Space for Transferring Knowledge

ECCV 2020poster

Knowledge distillation refers to the process of training a student network to achieve better accuracy by learning from a pre-trained teacher network. Most of the existing knowledge distillation methods direct the student to follow the teacher by matching the teacher's output, feature maps or their d…

Cited by 12SourcePDFScholar
2020

This Dataset Does Not Exist: Training Models from Generated Images

ICASSP 2020accepted

Current generative networks are increasingly proficient in generating high-resolution realistic images. These generative networks, especially the conditional ones, can potentially become a great tool for providing new image datasets. This naturally brings the question: Can we train a classifier only…

Cited by 0SourceScholar
2020

Toward Unsupervised, Multi-Object Discovery in Large-Scale Image Collections

ECCV 2020poster

multi-object discovery in large-scale image collections","This paper addresses the problem of discovering the objects present in a collection of images without any supervision. We build on the optimization approach of Vo {m et al.} [34] with several key novelties: (1) We propose a novel saliency-bas…

2019

Addressing Failure Prediction by Learning Model Confidence

NeurIPS 2019poster

Assessing reliably the confidence of a deep neural net and predicting its failures is of primary importance for the practical deployment of these models. In this paper, we propose a new target criterion for model confidence, corresponding to the True Class Probability (TCP). We show how using the TC…

2018

Finding Beans in Burgers: Deep Semantic-Visual Embedding With Localization

CVPR 2018poster

Several works have proposed to learn a two-path neural network that maps images and texts, respectively, to a same shared Euclidean space where geometry captures useful semantic relationships. Such a multi-modal embedding can be trained and used for various tasks, notably image captioning. In the pr…

2018

Self-Supervised Multi-Level Face Model Learning for Monocular Reconstruction at Over 250 Hz

CVPR 2018poster

The reconstruction of dense 3D models of face geometry and appearance from a single image is highly challenging and ill-posed. To constrain the problem, many approaches rely on strong priors, such as parametric face models learned from limited 3D scan data. However, prior models restrict generalizat…

Cited by 308SourcePDFScholar
2017

Informed source separation via compressive graph signal sampling

ICASSP 2017accepted

We propose a novel informed source separation method for audio object coding based on a recent sampling theory for smooth signals on graphs. Assuming that only one source is active at each time-frequency point, we compute an ideal map indicating which source is active at each time-frequency point at…

Cited by 0SourceScholar
2017

Motion informed audio source separation

ICASSP 2017accepted

In this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is appli…

Cited by 0SourceScholar
2016

Automatic allocation of NTF components for user-guided audio source separation

ICASSP 2016accepted

Nonnegative matrix or tensor factorization is a very popular approach for audio source separation. One important problem in nonnegative tensor factorization (NTF) in the context of user-guided audio source separation is the necessity to manually assign the NTF components to audio sources in order to…

Cited by 0SourceScholar
2016

Sketching for large-scale learning of mixture models

ICASSP 2016accepted

Learning parameters from voluminous data can be prohibitive in terms of memory and computational requirements. We propose a "compressive learning" framework where we first sketch the data by computing random generalized moments of the underlying probability distribution, then estimate mixture model…

Cited by 0SourceScholar
2015

Hybrid multi-layer deep CNN/aggregator feature for image classification

ICASSP 2015accepted

Deep Convolutional Neural Networks (DCNN) have established a remarkable performance benchmark in the field of image classification, displacing classical approaches based on hand-tailored aggregations of local descriptors. Yet DCNNs impose high computational burdens both at training and at testing ti…

Cited by 0SourceScholar
2015

Incremental dense multi-modal 3D scene reconstruction

IROS 2015poster

Aquiring reliable depth maps is an essential prerequisite for accurate and incremental 3D reconstruction used in a variety of robotics applications. Depth maps produced by affordable Kinect-like cameras have become a de-facto standard for indoor reconstruction and the driving force behind the succes…

Cited by 19SourceScholar
2015

Incremental dense semantic stereo fusion for large-scale semantic scene reconstruction

ICRA 2015poster

Our abilities in scene understanding, which allow us to perceive the 3D structure of our surroundings and intuitively recognise the objects we see, are things that we largely take for granted, but for robots, the task of understanding large scenes quickly remains extremely challenging. Recently, sce…

Cited by 260SourceScholar