← Search

Alexey Gritsenko

4 accepted papers

2024

Time- Memory- and Parameter-Efficient Visual Adaptation

CVPR 2024highlight

As foundation models become more popular there is a growing need to efficiently finetune them for downstream tasks. Although numerous adaptation methods have been proposed they are designed to be efficient only in terms of how many parameters are trained. They however typically still require backpro…

Cited by 15SourcePDFScholar
2023

Video OWL-ViT: Temporally-consistent Open-world Localization in Video

ICCV 2023poster

We present an architecture and a training recipe that adapts pretrained open-world image models to localization in videos. Understanding the open visual world (without being constrained by fixed label spaces) is crucial for many real-world vision tasks. Contrastive pre-training on large image-text d…

Cited by 17PDFScholar
2022

Simple Open-Vocabulary Object Detection with Vision Transformers

ECCV 2022poster

"Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are less well established, especially in the long-tailed and open-vocabulary setting, where training data is relatively sca…

2020

A Spectral Energy Distance for Parallel Speech Synthesis

NeurIPS 2020poster

Speech synthesis is an important practical generative modeling problem that has seen great progress over the last few years, with likelihood-based autoregressive neural models now outperforming traditional concatenative systems. A downside of such autoregressive models is that they require executing…