← Search

Julieta Martinez

12 accepted papers

2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

CVPR 2026

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and

Cited by 0SourcecodeScholar
2025

Pippo: High-Resolution Multi-View Humans from a Single Image

CVPR 2025highlight

We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs - e.g., a fitted parametric model or camera parameters of the input…

Cited by 1SourcePDFScholar
2024

Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars

NeurIPS 2024poster

To build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (_i.e._, easily adaptable to novel identities). Towards these goals, _paired_ captures, that…

2024

Sapiens: Foundation for Human Vision Models

ECCV 2024oral

"We present Sapiens, a family of models for four fundamental human-centric vision tasks – 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction. Our models natively support 1K high-resolution inference and are extremely easy to adapt for individual tasks by simp…

Cited by 28SourcePDFScholar
2021

Deep Multi-Task Learning for Joint Localization, Perception, and Prediction

CVPR 2021poster

Over the last few years, we have witnessed tremendous progress on many subtasks of autonomous driving including perception, motion forecasting, and motion planning. However, these systems often assume that the car is accurately localized against a high-definition map. In this paper we question this…

Cited by 46PDFScholar
2021

Permute, Quantize, and Fine-Tune: Efficient Compression of Neural Networks

CVPR 2021poster

Compressing large neural networks is an important step for their deployment in resource-constrained computational platforms. In this context, vector quantization is an appealing framework that expresses multiple parameters using a single code, and has recently achieved state-of-the-art network compr…

Cited by 49PDFcodeScholar
2020

Pit30M: A Benchmark for Global Localization in the Age of Self-Driving Cars

IROS 2020poster

We are interested in understanding whether retrieval-based localization approaches are good enough in the context of self-driving vehicles. Towards this goal, we introduce Pit30M, a new image and LiDAR dataset with over 30 million frames, which is 10 to 100 times larger than those used in previous w…

Cited by 15SourcecodeScholar
2019

Learning to Localize Through Compressed Binary Maps

CVPR 2019poster

One of the main difficulties of scaling current localization systems to large environments is the on-board storage required for the maps. In this paper we propose to learn to compress the map representation such that it is optimal for the localization task. As a consequence, higher compression rates…

Cited by 39PDFScholar
2018

LSQ++: Lower running time and higher recall in multi-codebook quantization

ECCV 2018poster

Multi-codebook quantization (MCQ) is the task of expressing a set of vectors as accurately as possible in terms of discrete entries in multiple bases. Work in MCQ is heavily focused on lowering quantization error, thereby improving distance estimation and recall on benchmarks of visual descriptors a…

2017

A Simple yet Effective Baseline for 3D Human Pose Estimation

ICCV 2017poster

Following the success of deep convolutional networks, state-of-the-art methods for 3d human pose estimation have focused on deep end-to-end systems that predict 3d joint locations given raw image pixels. Despite their excellent performance, it is often not easy to understand whether their remaining…

Cited by 1521PDFcodeScholar