← Search

Loris Bazzani

9 accepted papers

2025

Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval

ICCV 2025poster

Structuring latent representations in a hierarchical manner enables models to learn patterns at multiple levels of abstraction. However, most prevalent image understanding models focus on visual similarity, and learning visual hierarchies is relatively unexplored. In this work, for the first time, w…

Cited by 0SourcePDFScholar
2024

ViewFusion: Towards Multi-View Consistency via Interpolated Denoising

CVPR 2024poster

Novel-view synthesis through diffusion models has demonstrated remarkable potential for generating diverse and high-quality images. Yet the independent process of image generation in these prevailing methods leads to challenges in maintaining multiple-view consistency. To address this we introduce V…

2021

Learning Attribute-Driven Disentangled Representations for Interactive Fashion Retrieval

ICCV 2021poster

Interactive retrieval for online fashion shopping provides the ability of changing image retrieval results according to the user feedback. One common problem in interactive retrieval is that a specific user interaction (e.g., changing the color of a T-shirt) causes other aspects to change inadverten…

Cited by 58PDFcodeScholar
2021

Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised Learning

CVPR 2021poster

Cross-modal recipe retrieval has recently gained substantial attention due to the importance of food in people's lives, as well as the availability of vast amounts of digital cooking recipes and food images to train machine learning models. In this work, we revisit existing approaches for cross-moda…

Cited by 88PDFcodeScholar
2017

Recurrent Mixture Density Network for Spatiotemporal Visual Attention

ICLR 2017poster

In many computer vision tasks, the relevant information to solve the problem at hand is mixed to irrelevant, distracting information. This has motivated researchers to design attentional models that can dynamically focus on parts of images or videos that are salient, e.g., by down-weighting irreleva…

Cited by 168SourceScholar
2016

Approximate Log-Hilbert-Schmidt Distances Between Covariance Operators for Image Classification

CVPR 2016poster

This paper presents a novel framework for visual object recognition using infinite-dimensional covariance operators of input features, in the paradigm of kernel methods on infinite-dimensional Riemannian manifolds. Our formulation provides a rich representation of image features by exploiting their…

Cited by 24PDFScholar