← Search

Alessandro Gnutti

4 accepted papers

2025

Bridging Compressed Image Latents and Multimodal Large Language Models

ICLR 2025poster

This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hi…

Cited by 1SourcePDFScholar
2025

MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding

ICCV 2025poster

This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and short-term reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional video coding offer superior coding performance. However, the…

2022

CANF-VC: Conditional Augmented Normalizing Flows for Video Compression

ECCV 2022poster

"This paper presents an end-to-end learning-based video compression system, termed CANF-VC, based on conditional augmented normalizing flows (CANF). Most learned video compression systems adopt the same hybrid-based coding architecture as the traditional codecs. Recent research on conditional coding…