← Search

Adrian Bulat

32 accepted papers

2026

Hierarchical Image Tokenization for Multi-Scale Image Super Resolution

ICML 2026poster

We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. Recently, VAR models challenged the dominance of diffusion-based models by adopting a next-scale prediction paradigm. Specifically, VAR models iteratively estimate the…

Cited by 0SourceScholar
2026

Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration

CVPR 2026

Image restoration aims to recover high quality images from inputs degraded by various factors, such as adverse weather, blur, or low light. While recent studies have shown remarkable progress across individual or unified restoration tasks, they still suffer from limited generalization and inefficien

Cited by 0SourceScholar
2026

VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

CVPR 2026

Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on challenging tasks that require fine-grained understan

Cited by 0SourceScholar
2025

Compress & Cache: Vision token compression for efficient generation and retrieval

NeurIPS 2025poster

This work aims to compress the vision tokens of an LVLM into a representation that is simultaneously suitable for (a) generative and (b) discriminative tasks, (c) is nearly lossless, and (d) storage-efficient. To this end, we propose C&C, a novel compression method that leverages the LVLM itself fo…

Cited by 0SourceScholar
2025

FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion

CVPR 2025poster

Diffusion models are proficient at generating high-quality images. They are however effective only when operating at the resolution used during training. Inference at a scaled resolution leads to repetitive patterns and structural distortions. Retraining at higher resolutions quickly becomes prohib…

Cited by 1SourcePDFScholar
2025

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

EMNLP 2025

Contrastively-trained Vision-Language Models (VLMs), such as CLIP, have become the standard approach for learning discriminative vision-language representations. However, these models often exhibit shallow language understanding, manifesting bag-of-words behaviour. These limitations are reinforced b

Cited by 0SourcePDFScholar
2025

VladVA: Discriminative Fine-tuning of LVLMs

CVPR 2025poster

Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models have limited language understanding, often exhibiting a "bag of words" behavior. At the same time, Large Vision-Language M…

Cited by 0SourcePDFScholar
2024

Efficient Vision-Language pre-training via domain-specific learning for human activities

EMNLP 2024main

Current Vision-Language (VL) models owe their success to large-scale pre-training on web-collected data, which in turn requires high-capacity architectures and large compute resources for training. We posit that when the downstream tasks are known in advance, which is in practice common, the pretrai…

2024

FFF: Fixing Flawed Foundations in Contrastive Pre-Training Results in Very Strong Vision-Language Models

CVPR 2024poster

Despite noise and caption quality having been acknowledged as important factors impacting vision-language contrastive pre-training in this paper we show that the full potential of improving the training process by addressing such issues is yet to be realized. Specifically we firstly study and analyz…

Cited by 5SourcePDFScholar
2023

Bayesian Prompt Learning for Image-Language Model Generalization

ICCV 2023poster

Foundational image-language models have generated considerable interest due to their efficient adaptation to downstream tasks by prompt learning. Prompt learning treats part of the language model input as trainable while freezing the rest, and optimizes an Empirical Risk Minimization objective. Howe…

Cited by 42PDFcodeScholar
2023

Black Box Few-Shot Adaptation for Vision-Language Models

ICCV 2023poster

Vision-Language (V-L) models trained with contrastive learning to align the visual and language modalities have been shown to be strong few-shot learners. Soft prompt learning is the method of choice for few-shot downstream adaption aiming to bridge the modality gap caused by the distribution shift…

Cited by 47PDFcodeScholar
2023

FS-DETR: Few-Shot DEtection TRansformer with Prompting and without Re-Training

ICCV 2023poster

This paper is on Few-Shot Object Detection (FSOD), where given a few templates (examples) depicting a novel class (not seen during training), the goal is to detect all of its occurrences within a set of images. From a practical perspective, an FSOD system must fulfil the following desiderata: (a) it…

Cited by 42PDFScholar
2023

LASP: Text-to-Text Optimization for Language-Aware Soft Prompting of Vision & Language Models

CVPR 2023poster

Soft prompt learning has recently emerged as one of the methods of choice for adapting V&L models to a downstream task using a few training examples. However, current methods significantly overfit the training data, suffering from large accuracy degradation when tested on unseen classes from the sam…

2023

ReGen: A good Generative Zero-Shot Video Classifier Should be Rewarded

ICCV 2023poster

This paper sets out to solve the following problem: How can we turn a generative video captioning model into an open-world video/action classification model? Video captioning models can naturally produce open-ended free-form descriptions of a given video which, however, might not be discriminative e…

Cited by 2PDFScholar
2022

EdgeViTs: Competing Light-Weight CNNs on Mobile Devices with Vision Transformers

ECCV 2022poster

"Self-attention based models such as vision transformers (ViTs) have emerged as a very competitive architecture alternative to convolutional neural networks (CNNs) in computer vision. Despite increasingly stronger variants with ever-higher recognition accuracies, due to the quadratic complexity of s…

2022

Pre-training Strategies and Datasets for Facial Representation Learning

ECCV 2022poster

"What is the best way to learn a universal face representation? Recent work on Deep Learning in the area of face analysis has focused on supervised learning for specific tasks of interest (e.g. face recognition, facial landmark localization etc.) but has overlooked the overarching question of how to…

2021

Improving Memory Banks for Unsupervised Learning with Large Mini-Batch, Consistency and Hard Negative Mining

ICASSP 2021accepted

An important component of unsupervised learning by instance-based discrimination is a memory bank for storing a feature representation for each training sample in the dataset. In this paper, we introduce 3 improvements to the vanilla memory bank-based formulation which brings massive accuracy gains:…

Cited by 0SourceScholar
2021

Knowledge distillation via softmax regression representation learning

ICLR 2021poster

This paper addresses the problem of model compression via knowledge distillation. We advocate for a method that optimizes the output feature of the penultimate layer of the student network and hence is directly related to representation learning. Previous distillation methods which typically impose…

2021

Space-time Mixing Attention for Video Transformer

NeurIPS 2021poster

This paper is on video recognition using Transformers. Very recent attempts in this area have demonstrated promising results in terms of recognition accuracy, yet they have been also shown to induce, in many cases, significant computational overheads due to the additional modelling of the temporal i…

2020

Factorized Higher-Order CNNs With an Application to Spatio-Temporal Emotion Estimation

CVPR 2020poster

Training deep neural networks with spatio-temporal (i.e., 3D) or multidimensional convolutions of higher-order is computationally challenging due to millions of unknown parameters across dozens of layers. To alleviate this, one approach is to apply low-rank tensor decompositions to convolution kerne…

Cited by 109PDFScholar
2020

Towards Pose-Invariant Lip-Reading

ICASSP 2020accepted

Lip-reading models have been significantly improved recently thanks to powerful deep learning architectures. However, most works focused on frontal or near frontal views of the mouth. As a consequence, lip-reading performance seriously deteriorates in non-frontal mouth views. In this work, we presen…

Cited by 0SourceScholar
2020

Training binary neural networks with real-to-binary convolutions

ICLR 2020poster

This paper shows how to train binary networks to within a few percent points (~3-5%) of the full precision counterpart. We first show how to build a strong baseline, which already achieves state-of-the-art accuracy, by combining recently proposed advances and carefully adjusting the optimization pro…

Cited by 298SourcecodeScholar
2019

T-Net: Parametrizing Fully Convolutional Nets With a Single High-Order Tensor

CVPR 2019poster

Recent findings indicate that over-parametrization, while crucial for successfully training deep neural networks, also introduces large amounts of redundancy. Tensor methods have the potential to efficiently parametrize over-complete representations by leveraging this redundancy. In this paper, we p…

Cited by 93PDFScholar
2018

Super-FAN: Integrated Facial Landmark Localization and Super-Resolution of Real-World Low Resolution Faces in Arbitrary Poses With GANs

CVPR 2018poster

This paper addresses 2 challenging tasks: improving the quality of low resolution facial images and accurately locating the facial landmarks on such poor resolution images. To this end, we make the following 5 contributions: (a) we propose Super-FAN: the very first end-to-end system that addresses b…

2018

To learn image super-resolution, use a GAN to learn how to do image degradation first

ECCV 2018poster

This paper is on image and face super-resolution. The vast majority of prior work for this problem focus on how to increase the resolution of low-resolution images which are artificially generated by simple bilinear down-sampling (or in a few cases by blurring followed by down-sampling). We show tha…

2017

Binarized Convolutional Landmark Localizers for Human Pose Estimation and Face Alignment With Limited Resources

ICCV 2017oral

Our goal is to design architectures that retain the groundbreaking performance of CNNs for landmark localization and at the same time are lightweight, compact and suitable for applications with limited computational resources. To this end, we make the following contributions: (a) we are the first to…

Cited by 279PDFcodeScholar
2017

How Far Are We From Solving the 2D & 3D Face Alignment Problem? (And a Dataset of 230,000 3D Facial Landmarks)

ICCV 2017poster

This paper investigates how far a very deep neural network is from attaining close to saturating performance on existing 2D and 3D face alignment datasets. To this end, we make the following 5 contributions: (a) we construct, for the first time, a very strong baseline by combining a state-of-the-art…

Cited by 1952PDFcodeScholar
2017

Large Pose 3D Face Reconstruction From a Single Image via Direct Volumetric CNN Regression

ICCV 2017poster

3D face reconstruction is a fundamental Computer Vision problem of extraordinary difficulty. Current systems often assume the availability of multiple facial images (sometimes from the same subject) as input, and must address a number of methodological challenges such as establishing dense correspon…

Cited by 579PDFcodeScholar