← Search

Jean Ponce

45 accepted papers

2026

On the Conic Complementarity of Planar Contacts

ICRA 2026poster

We present a unifying theoretical result that con- nects two foundational principles in robotics: the Signorini law for point contacts, which underpins many simulation methods for preventing object interpenetration, and the center of pres- sure (also known as the zero-moment point), a key concept in…

2026

Optimal transport unlocks end-to-end learning for single-molecule localization

ICLR 2026poster

Single‑molecule localization microscopy (SMLM) allows reconstructing cellular organelles and biology-relevant structures far beyond the limited spatial resolution imposed by optics constrains, using tagged biomolecule positions. Currently, efficient SMLM requires non‑overlapping emitting fluorophore…

Cited by 0SourceScholar
2025

A New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging Observations

CVPR 2025poster

The search for exoplanets is an active field in astronomy, with direct imaging as one of the most challenging methods due to faint exoplanet signals buried within stronger residual starlight. Successful detection requires advanced image processing to separate the exoplanet signal from this nuisance…

2023

An Image Quality Assessment Dataset for Portraits

CVPR 2023poster

Year after year, the demand for ever-better smartphone photos continues to grow, in particular in the domain of portrait photography. Manufacturers thus use perceptual quality criteria throughout the development of smartphone cameras. This costly procedure can be partially replaced by automated lear…

2023

Learning Reward Functions for Robotic Manipulation by Observing Humans

ICRA 2023poster

Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not least a difference in action and observation spaces. In this…

Cited by 26SourceScholar
2023

WALDO: Future Video Synthesis Using Object Layer Decomposition and Parametric Flow Prediction

ICCV 2023poster

This paper presents WALDO (WArping Layer-Decomposed Objects), a novel approach to the prediction of future video frames from past ones. Individual images are decomposed into multiple layers combining object masks and a small set of control points. The layer structure is shared across all frames in e…

Cited by 8PDFcodeScholar
2022

Active Learning Strategies for Weakly-Supervised Object Detection

ECCV 2022poster

"Object detectors trained with weak annotations are affordable alternatives to fully-supervised counterparts. However, there is still a significant performance gap between them. We propose to narrow this gap by fine-tuning a base pre-trained weakly-supervised detector with a few fully-annotated samp…

2022

Assembly Planning from Observations under Physical Constraints

IROS 2022poster

This paper addresses the problem of copying an unknown assembly of primitives with known shape and appearance using information extracted from a single photograph by an off-the-shelf procedure for object detection and pose estimation. The proposed algorithm uses a simple combination of physical stab…

Cited by 5SourceScholar
2022

VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

ICLR 2022poster

Recent self-supervised methods for image representation learning maximize the agreement between embedding vectors produced by encoders fed with different views of the same image. The main challenge is to prevent a collapse in which the encoders produce constant or non-informative vectors. We introd…

2021

Large-Scale Unsupervised Object Discovery

NeurIPS 2021poster

Existing approaches to unsupervised object discovery (UOD) do not scale up to large datasets without approximations that compromise their performance. We propose a novel formulation of UOD as a ranking problem, amenable to the arsenal of distributed methods available for eigenvalue problems and link…

2021

Online Learning and Control of Complex Dynamical Systems from Sensory Input

NeurIPS 2021poster

Identifying an effective model of a dynamical system from sensory data and using it for future state prediction and control is challenging. Recent data-driven algorithms based on Koopman theory are a promising approach to this problem, but they typically never update the model once it has been ident…

2021

Unsupervised Layered Image Decomposition Into Object Prototypes

ICCV 2021poster

We present an unsupervised learning framework for decomposing images into layers of automatically discovered object models. Contrary to recent approaches that model image layers with autoencoder networks, we represent them as explicit transformations of a small set of prototypical images. Our model…

Cited by 65PDFcodeScholar
2020

A Flexible Framework for Designing Trainable Priors with Adaptive Smoothing and Game Encoding

NeurIPS 2020poster

We introduce a general framework for designing and training neural network layers whose forward passes can be interpreted as solving non-smooth convex optimization problems, and whose architectures are derived from an optimization algorithm. We focus on convex games, solved by local agents represent…

Cited by 21SourcePDFScholar
2020

Fully Trainable and Interpretable Non-Local Sparse Models for Image Restoration

ECCV 2020poster

Non-local self-similarity and sparsity principles have proven to be powerful priors for natural image modeling. We propose a novel differentiable relaxation of joint sparsity that exploits both principles and leads to a general framework for image restoration which is (1) trainable end to end, (2) f…

2020

Learning to Compose Hypercolumns for Visual Correspondence

ECCV 2020poster

Feature representation plays a crucial role in visual correspondence, and recent methods for image matching resort to deeply stacked convolutional layers. These models, however, are both monolithic and static in the sense that they typically use a specific level of features, e.g., the output of the…

2020

Minimal Solutions to Relative Pose Estimation From Two Views Sharing a Common Direction With Unknown Focal Length

CVPR 2020poster

We propose minimal solutions to relative pose estimation problem from two views sharing a common direction with unknown focal length. This is relevant for cameras equipped with an IMU (inertial measurement unit), e.g., smart phones, tablets. Similar to the 6-point algorithm for two cameras with unkn…

Cited by 25PDFScholar
2020

Toward Unsupervised, Multi-Object Discovery in Large-Scale Image Collections

ECCV 2020poster

multi-object discovery in large-scale image collections","This paper addresses the problem of discovering the objects present in a collection of images without any supervision. We build on the optimization approach of Vo {m et al.} [34] with several key novelties: (1) We propose a novel saliency-bas…

2019

An Efficient Solution to the Homography-Based Relative Pose Problem With a Common Reference Direction

ICCV 2019oral

In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with a common reference direction. We explore the rank-1 constraint on the difference between the Euclidean homography matrix and the corresponding rotation, and propose an efficient two-st…

Cited by 27PDFScholar
2019

Build your own hybrid thermal/EO camera for autonomous vehicle

ICRA 2019poster

In this work, we propose a novel paradigm to design a hybrid thermal/EO (Electro-Optical or visible-light) camera, whose thermal and RGB frames are pixel-wisely aligned and temporally synchronized. Compared with the existing schemes, we innovate in three ways in order to make it more compact in dime…

Cited by 8SourceScholar
2019

Coordinate-Free Carlsson-Weinshall Duality and Relative Multi-View Geometry

CVPR 2019oral

We present a coordinate-free description of Carlsson-Weinshall duality between scene points and camera pinholes and use it to derive a new characterization of primal/dual multi-view geometry. In the case of three views, a particular set of reduced trilinearities provide a novel parameterization of c…

Cited by 9PDFcodeScholar
2019

Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural Features

ICCV 2019poster

Establishing visual correspondences under large intra-class variations requires analyzing images at different levels, from features linked to semantics and context to local patterns, while being invariant to instance-specific details. To tackle these challenges, we represent images by "hyperpixels"…

Cited by 121PDFcodeScholar
2019

Unsupervised Image Matching and Object Discovery as Optimization

CVPR 2019poster

Learning with complete or partial supervision is power- ful but relies on ever-growing human annotation efforts. As a way to mitigate this serious problem, as well as to serve specific applications, unsupervised learning has emerged as an important field of research. In computer vision, unsu- pervis…

Cited by 77PDFcodeScholar
2018

Dijkstra Model for Stereo-Vision Based Road Detection: A Non-Parametric Method

ICRA 2018poster

This paper proposes a new method for detecting a road from a stereo pair of images. First, the horizon is accurately estimated by a robust, weighted-sampling RANSAC-like method in the improved v-disparity map. The vanishing point of the road region is located using both the horizon information and r…

Cited by 7SourceScholar
2017

General Models for Rational Cameras and the Case of Two-Slit Projections

CVPR 2017poster

The rational camera model recently introduced in [18] provides a general methodology for studying abstract nonlinear imaging systems and their multi-view geometry. This paper builds on this framework to study "physical realizations" of rational cameras. More precisely, we give an explicit account of…

Cited by 10PDFScholar
2017

Kernel Square-Loss Exemplar Machines for Image Retrieval

CVPR 2017poster

Zepeda and Perez have recently demonstrated the promise of the exemplar SVM (ESVM) as a feature encoder for image retrieval. This paper extends this approach in several directions: We first show that replacing the hinge loss by the square loss in the ESVM cost function significantly reduces encoding…

Cited by 13PDFScholar
2017

SCNet: Learning Semantic Correspondence

ICCV 2017poster

This paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearan…

Cited by 159PDFcodeScholar
2016

Proposal Flow

CVPR 2016poster

Finding image correspondences remains a challenging problem in the presence of intra-class variations and large changes in scene layout. Semantic flow methods are designed to handle images depicting different instances of the same object or scene category. We introduce a novel approach to semantic…

Cited by 164PDFScholar
2015

Learning a Convolutional Neural Network for Non-Uniform Motion Blur Removal

CVPR 2015poster

In this paper, we address the problem of estimating and removing non-uniform motion blur from a single blurry image. We propose a deep learning approach to predicting the probabilistic distribution of motion blur at the patch level using a convolutional neural network (CNN). We further extend the ca…

Cited by 1136SourcePDFScholar
2015

Unsupervised Object Discovery and Localization in the Wild: Part-Based Matching With Bottom-Up Region Proposals

CVPR 2015poster

This paper addresses unsupervised discovery and localization of dominant objects from a noisy image collection with multiple object classes. The setting of this problem is fully unsupervised, without even image-level annotations or any assumption of a single dominant class. This is far more general…

Cited by 321SourcePDFScholar
2015

Unsupervised Object Discovery and Tracking in Video Collections

ICCV 2015poster

This paper addresses the problem of automatically localizing dominant objects as spatio-temporal tubes in a noisy collection of videos with minimal or even no supervision. We formulate the problem as a combination of two complementary processes: discovery and tracking. The first one establishes corr…

Cited by 153PDFScholar
2015

Weakly-Supervised Alignment of Video With Text

ICCV 2015poster

Suppose that we are given a set of videos, along with natural language descriptions in the form of multiple sentences (e.g., manual annotations, movie scripts, sport summaries etc.), and that these sentences appear in the same temporal order as their visual counterparts. We propose in this paper a m…

Cited by 171PDFcodeScholar