← Search

Manuel Giollo

3 accepted papers

2025

Detecting and Mitigating Challenges in Zero-Shot Video Summarization with Video LLMs

ACL 2025finding

Video summarization aims to generate a condensed textual version of an original video. Summaries may consist of either plain text or a shortlist of salient events, possibly including temporal or spatial references. Video Large Language Models (VLLMs) exhibit impressive zero-shot capabilities in vide…

2023

Exploring Subgroup Performance in End-to-End Speech Models

ICASSP 2023accepted

End-to-End Spoken Language Understanding models are generally evaluated according to their overall accuracy, or separately on (a priori defined) data subgroups of interest. We propose a technique for analyzing model performance at the subgroup level, which considers all subgroups that can be defined…

Cited by 0SourceScholar
2021

Improved Robustness to Disfluencies in Rnn-Transducer Based Speech Recognition

ICASSP 2021accepted

Automatic Speech Recognition (ASR) based on Recurrent Neural Network Transducers (RNN-T) is gaining interest in the speech community. We investigate data selection and preparation choices aiming for improved robustness of RNN-T ASR to speech disfluencies with a focus on partial words. For evaluation…

Cited by 0SourceScholar