← Search

Alessio Brutti

13 accepted papers

2026

DISTILLATION-BASED LAYER DROPPING (DLD): EFFECTIVE END-TO-END FRAMEWORK FOR DYNAMIC SPEECH NETWORKS

ICASSP 2026poster

Edge devices operate in constrained and varying resource settings, requiring dynamic architectures that can adapt to limitations of the available resources. To meet such demands, layer dropping ($\mathcal{LD}$) approach is typically used to transform static models into dynamic ones by skipping parts…

Cited by 0SourcePDFScholar
2025

EFL-PEFT: A communication Efficient Federated Learning framework using PEFT sparsification for ASR

ICASSP 2025accepted

Federated Learning (FL) has garnered substantial interest in training different speech-based tasks (e.g. automatic speech recognition (ASR), and other speech classification tasks): recently, fine-tuning pre-trained self-supervised models for different speech-based tasks has shown promising performan…

Cited by 0SourceScholar
2025

Large Language Models are Strong Audio-Visual Speech Recognition Learners

ICASSP 2025accepted

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the au…

Cited by 0SourceScholar
2024

Continual Contrastive Spoken Language Understanding

ACL 2024findings

Recently, neural networks have shown impressive progress across diverse fields, with speech processing being no exception. However, recent breakthroughs in this area require extensive offline training using large datasets and tremendous computing resources. Unfortunately, these models struggle to re…

Cited by 2SourcePDFScholar
2024

MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages

EMNLP 2024main

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing speech FMs (SFMs) fall short of full compliance with the open-source principles, even if claimed otherwise, as no existin…

2022

End-to-End Low Resource Keyword Spotting Through Character Recognition and Beam-Search Re-Scoring

ICASSP 2022accepted

This paper describes an end-to-end approach to perform keyword spotting with a pre-trained acoustic model that uses recurrent neural networks and connectionist temporal classification loss. Our approach is specifically designed for low-resource keyword spotting tasks where extremely small amounts of…

Cited by 0SourceScholar
2022

Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?

ICASSP 2022accepted

Humans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of modalities. Generally, models that fuse complementary information…

Cited by 0SourceScholar
2022

Scalable Neural Architectures for End-to-End Environmental Sound Classification

ICASSP 2022accepted

Sound Event Detection (SED) is a complex task simulating human ability to recognize what is happening in the surrounding from auditory signals only. This technology is a crucial asset in many applications such as smart cities. Here, urban sounds can be detected and processed by embedded devices in a…

Cited by 0SourceScholar
2021

Robust Latent Representations Via Cross-Modal Translation and Alignment

ICASSP 2021accepted

Multi-modal learning relates information across observation modalities of the same physical phenomenon to leverage complementary information. Most multi-modal machine learning methods require that all the modalities used for training are also available for testing. This is a limitation when signals…

Cited by 0SourceScholar
2019

Accurate Target Annotation in 3D from Multimodal Streams

ICASSP 2019accepted

Accurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverag…

Cited by 0SourceScholar
2018

3D Mouth Tracking from a Compact Microphone Array Co-Located with a camera

ICASSP 2018accepted

We address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood c…

Cited by 0SourceScholar
2017

3D audio-visual speaker tracking with an adaptive particle filter

ICASSP 2017accepted

We propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter frame…

Cited by 0SourceScholar