← Search

Enrique Sanchez

9 accepted papers

2026

Hierarchical Image Tokenization for Multi-Scale Image Super Resolution

ICML 2026poster

We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. Recently, VAR models challenged the dominance of diffusion-based models by adopting a next-scale prediction paradigm. Specifically, VAR models iteratively estimate the…

Cited by 0SourceScholar
2026

Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration

CVPR 2026

Image restoration aims to recover high quality images from inputs degraded by various factors, such as adverse weather, blur, or low light. While recent studies have shown remarkable progress across individual or unified restoration tasks, they still suffer from limited generalization and inefficien

Cited by 0SourceScholar
2024

Multiscale Vision Transformers Meet Bipartite Matching for Efficient Single-stage Action Localization

CVPR 2024poster

Action Localization is a challenging problem that combines detection and recognition tasks which are often addressed separately. State-of-the-art methods rely on off-the-shelf bounding box detections pre-computed at high resolution and propose transformer models that focus on the classification task…

2023

Bayesian Prompt Learning for Image-Language Model Generalization

ICCV 2023poster

Foundational image-language models have generated considerable interest due to their efficient adaptation to downstream tasks by prompt learning. Prompt learning treats part of the language model input as trainable while freezing the rest, and optimizes an Empirical Risk Minimization objective. Howe…

Cited by 42PDFcodeScholar
2023

ReGen: A good Generative Zero-Shot Video Classifier Should be Rewarded

ICCV 2023poster

This paper sets out to solve the following problem: How can we turn a generative video captioning model into an open-world video/action classification model? Video captioning models can naturally produce open-ended free-form descriptions of a given video which, however, might not be discriminative e…

Cited by 2PDFScholar
2022

Pre-training Strategies and Datasets for Facial Representation Learning

ECCV 2022poster

"What is the best way to learn a universal face representation? Recent work on Deep Learning in the area of face analysis has focused on supervised learning for specific tasks of interest (e.g. face recognition, facial landmark localization etc.) but has overlooked the overarching question of how to…

2021

Affective Processes: Stochastic Modelling of Temporal Context for Emotion and Facial Expression Recognition

CVPR 2021poster

Temporal context is key to the recognition of expressions of emotion. Existing methods, that rely on recurrent or self-attention models to enforce temporal consistency, work on the feature level, ignoring the task-specific temporal dependencies, and fail to model context uncertainty. To alleviate th…

Cited by 42PDFcodeScholar
2020

Unsupervised Learning of Object Landmarks via Self-Training Correspondence

NeurIPS 2020poster

This paper addresses the problem of unsupervised discovery of object landmarks. We take a different path compared to that of existing works, based on 2 novel perspectives: (1) Self-training: starting from generic keypoints, we propose a self-training approach where the goal is to learn a detector th…