← Search

Romain Brégier

9 accepted papers

2025

Disentangled Object-Centric Image Representation for Robotic Manipulation

IROS 2025

Learning robotic manipulation skills from vision is a promising approach for developing robotics applications that can generalize broadly to real-world scenarios. As such, many approaches to enable this vision have been explored with fruitful results. Particularly, object-centric representation meth

Cited by 2SourceScholar
2024

Cross-view and Cross-pose Completion for 3D Human Understanding

CVPR 2024poster

Human perception and understanding is a major domain of computer vision which like many other vision subdomains recently stands to gain from the use of large models pre-trained on large datasets. We hypothesize that the most common pre-training strategy of relying on general purpose object-centric i…

Cited by 5SourcePDFScholar
2024

MFOS: Model-Free & One-Shot Object Pose Estimation

AAAI 2024technical

Existing learning-based methods for object pose estimation in RGB images are mostly model-specific or category based. They lack the capability to generalize to new object categories at test time, hence severely hindering their practicability and scalability. Notably, recent attempts have been made t…

Cited by 4SourcePDFScholar
2024

Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot

ECCV 2024poster

"We present , a strong model for multi-person 3D human mesh recovery from a single RGB image. Predictions encompass the whole body, , including hands and facial expressions, using the SMPL-X parametric model and 3D location in the camera coordinate system. Our model detects people by predicting coar…

2023

CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow

ICCV 2023poster

Despite impressive performance for high-level downstream tasks, self-supervised pre-training methods have not yet fully delivered on dense geometric vision tasks such as stereo matching or optical flow. The application of self-supervised concepts, such as instance discrimination or masked image mode…

Cited by 100PDFcodeScholar
2022

CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

NeurIPS 2022accept

Masked Image Modeling (MIM) has recently been established as a potent pre-training paradigm. A pretext task is constructed by masking patches in an input image, and this masked content is then predicted by a neural network using visible patches as sole input. This pre-training leads to state-of-the-…

2020

DOPE: Distillation Of Part Experts for whole-body 3D pose estimation in the wild

ECCV 2020poster

We introduce DOPE, the first method to detect and estimate whole-body 3D human poses, including bodies, hands and faces, in the wild. Achieving this level of details is key for a number of applications that require understanding the interactions of the people with each other or with the environment.…

Cited by 65SourcePDFScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…