← Search

Filip Radenovic

14 accepted papers

2024

Context Diffusion: In-Context Aware Image Generation

ECCV 2024poster

"We propose Context Diffusion, a diffusion-based framework that enables image generation models to learn from visual examples presented in context. Recent work tackles such in-context learning for image generation, where a query image is provided alongside context examples and text prompts. However,…

Cited by 9SourcePDFScholar
2023

Cola: A Benchmark for Compositional Text-to-image Retrieval

NeurIPS 2023poster

Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combining objects with their attributes. To measure this lack of compositional capability, we design Cola, a text-to-image retr…

Cited by 36SourcePDFScholar
2023

Filtering, Distillation, and Hard Negatives for Vision-Language Pre-Training

CVPR 2023poster

Vision-language models trained with contrastive learning on large-scale noisy data are becoming increasingly popular for zero-shot recognition problems. In this paper we improve the following three aspects of the contrastive pre-training pipeline: dataset noise, model initialization and the training…

2022

Making Heads or Tails: Towards Semantically Consistent Visual Counterfactuals

ECCV 2022poster

"A visual counterfactual explanation replaces image regions in a query image with regions from a distractor image such that the system’s decision on the transformed image changes to the distractor class. In this work, we present a novel framework for computing visual counterfactual explanations base…

2019

Targeted Mismatch Adversarial Attack: Query With a Flower to Retrieve the Tower

ICCV 2019poster

Access to online visual search engines implies sharing of private user content -- the query images. We introduce the concept of targeted mismatch attack for deep learning based retrieval systems to generate an adversarial image to conceal the query image. The generated image looks nothing like the u…

Cited by 79PDFcodeScholar
2018

Repeatability Is Not Enough: Learning Affine Regions via Discriminability

ECCV 2018poster

A method for learning local affine-covariant regions is presented. We show that maximizing geometric repeatability does not lead to local regions, a.k.a features, that are reliably matched and this necessitates descriptor-based learning. We explore factors that influence such learning and registrati…

2017

Working hard to know your neighbor's margins: Local descriptor learning loss

NeurIPS 2017poster

We introduce a loss for metric learning, which is inspired by the Lowe's matching criterion for SIFT. We show that the proposed loss, that maximizes the distance between the closest positive and closest negative example in the batch, is better than complex regularization methods; it works well for b…

2016

From Dusk Till Dawn: Modeling in the Dark

CVPR 2016spotlight

Internet photo collections naturally contain a large variety of illumination conditions, with the largest difference between day and night images. Current modeling techniques do not embrace the broad illumination range often leading to reconstruction failure or severe artifacts. We present an algori…

Cited by 52PDFScholar
2015

From Single Image Query to Detailed 3D Reconstruction

CVPR 2015poster

Structure-from-Motion for unordered image collections has significantly advanced in scale over the last decade. This impressive progress can be in part attributed to the introduction of efficient retrieval methods for those systems. While this boosts scalability, it also limits the amount of detail…

Cited by 144SourcePDFScholar
2015

Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition

ICCV 2015poster

Structure-based localization is the task of finding the absolute pose of a given query image w.r.t. a pre-computed 3D model. While this is almost trivial at small scale, special care must be taken as the size of the 3D model grows, because straight-forward descriptor matching becomes ineffective due…

Cited by 200PDFScholar