← Search

Lorenzo Vaiani

2 accepted papers

2025

Detecting and Mitigating Challenges in Zero-Shot Video Summarization with Video LLMs

ACL 2025finding

Video summarization aims to generate a condensed textual version of an original video. Summaries may consist of either plain text or a shortlist of salient events, possibly including temporal or spatial references. Video Large Language Models (VLLMs) exhibit impressive zero-shot capabilities in vide…

2024

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

ACL 2024findings

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation betwe…