← Search

Fabio Antonacci

23 accepted papers

2026

TRAINING-FREE MULTIMODAL GUIDANCE FOR VIDEO TO AUDIO GENERATION

ICASSP 2026poster

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results, existing approaches either require costly joint training on larg…

Cited by 0SourcePDFScholar
2025

A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction

ICASSP 2025accepted

Sound field reconstruction aims to estimate pressure fields in areas lacking direct measurements. Existing techniques often rely on strong assumptions or face challenges related to data availability or the explicit modeling of physical properties. To bridge these gaps, this study introduces a zero-s…

Cited by 0SourceScholar
2025

MambaFoley: Foley Sound Generation using Selective State-Space Models

ICASSP 2025accepted

Recent advancements in deep learning have led to widespread use of techniques for audio content generation, notably employing Denoising Diffusion Probabilistic Models (DDPM) across various tasks. Among these, Foley Sound Synthesis is of particular interest for its role in applications for the creati…

Cited by 0SourceScholar
2025

Past, Present, and Future of Spatial Audio and Room Acoustics

ICASSP 2025accepted

The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have been developed based both on theoretical advancements and pract…

Cited by 0SourceScholar
2025

Towards HRTF Personalization using Denoising Diffusion Models

ICASSP 2025accepted

Head-Related Transfer Functions (HRTFs) have fundamental applications for realistic rendering in immersive audio scenarios. However, they are strongly subject-dependent as they vary considerably depending on the shape of the ears, head and torso. Thus, personalization procedures are required for acc…

Cited by 0SourceScholar
2024

Reconstruction of Sound Field Through Diffusion Models

ICASSP 2024accepted

Reconstructing the sound field in a room is an important task for several applications, such as sound control and augmented (AR) or virtual reality (VR). In this paper, we propose a data-driven generative model for reconstructing the magnitude of acoustic fields in rooms with a focus on the modal fr…

Cited by 0SourceScholar
2023

Acoustic Source Localization in the Spherical Harmonics Domain Exploiting Low-Rank Approximations

ICASSP 2023accepted

Acoustic signal processing in the spherical harmonics domain (SHD) is an active research area that exploits the signals acquired by higher order microphone arrays. A very important task is that concerning the localization of active sound sources. In this paper, we propose a simple yet effective meth…

Cited by 0SourceScholar
2023

Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural Network

ICASSP 2023accepted

The interpretation and explanation of decision-making processes of neural networks are becoming a key factor in the deep learning field. Although several approaches have been presented for classification problems, the application to regression models needs to be further investigated. In this manuscr…

Cited by 0SourceScholar
2023

Real-Time Multichannel Speech Separation and Enhancement Using a Beamspace-Domain-Based Lightweight CNN

ICASSP 2023accepted

The problems of speech separation and enhancement concern the extraction of the speech emitted by a target speaker when placed in a scenario where multiple interfering speakers or noise are present, respectively. A plethora of practical applications such as home assistants and teleconferencing requi…

Cited by 0SourceScholar
2023

Zero-Shot Anomalous Sound Detection in Domestic Environments Using Large-Scale Pretrained Audio Pattern Recognition Models

ICASSP 2023accepted

Anomalous sound detection is central to audio-based surveillance and monitoring. In a domestic environment, however, the classes of sounds to be considered anomalous are situation-dependent and cannot be determined in advance. At the same time, it is not feasible to expect a demanding labeling effor…

Cited by 7SourceScholar
2022

A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording

ICASSP 2022accepted

Speech audio acquisitions exhibit different quality and reverberation properties depending on the recording setup and environment. For this reason, it is expected that speech analysis systems that work correctly on certain audio recordings may fail on others acquired in different acoustic contexts.…

Cited by 0SourceScholar
2022

Deepfake Speech Detection Through Emotion Recognition: A Semantic Approach

ICASSP 2022accepted

In recent years, audio and video deepfake technology has advanced relentlessly, severely impacting people’s reputation and reliability. Several factors have facilitated the growing deepfake threat. On the one hand, the hyper-connected society of social and mass media enables the spread of multimedia…

Cited by 0SourceScholar
2022

On the Prediction of the Frequency Response of a Wooden Plate from Its Mechanical Parameters

ICASSP 2022accepted

Inspired by deep learning applications in structural mechanics, we focus on how to train two predictors to model the relation between the vibrational response of a prescribed point of a wooden plate and its material properties. In particular, the eigenfrequencies of the plate are estimated via multi…

Cited by 0SourceScholar
2022

Sparsity-Based Sound Field Separation in the Spherical Harmonics Domain

ICASSP 2022accepted

Sound field analysis and reconstruction has been a topic of intense research in the last decades for its multiple applications in spatial audio processing tasks. In this context, the identification of the direct and reverberant sound field components is a problem of great interest, where several sol…

Cited by 27SourceScholar
2021

Arrays of First-Order Steerable Differential Microphones

ICASSP 2021accepted

The literature is rich with techniques for the design of small-size Differential Microphone Arrays (DMAs), known for their almost frequency-invariant beampatterns and low computational cost. Few works, instead, discuss the properties of beamformers based on multiple DMA units. In this paper, we cons…

Cited by 0SourceScholar
2021

Interpolation of Irregularly Sampled Frequency Response Functions Using Convolutional Neural Networks

ICASSP 2021accepted

In the field of structural mechanics, classical methods for the vibrational characterization of objects exploit the inherent redundancy of a relevant amount of measurements acquired over regular sampling grids. However, there are cases in which parts of the objects under analysis are not accessible…

Cited by 0SourceScholar
2020

Time Difference of Arrival Estimation from Frequency-Sliding Generalized Cross-Correlations Using Convolutional Neural Networks

ICASSP 2020accepted

The interest in deep learning methods for solving traditional signal processing tasks has been steadily growing in the last years. Time delay estimation (TDE) in adverse scenarios is a challenging problem, where classical approaches based on generalized cross-correlations (GCCs) have been widely use…

Cited by 0SourceScholar
2018

A Weighted Least Squares Beam Shaping Technique for Sound Field Control

ICASSP 2018accepted

A weighted least squares beam shaping technique for sound field control using a loudspeaker array is proposed. Given a desired spatial response at prescribed control points, the space-time filter is designed by solving a least squares minimization problem. To reduce the computational effort, we prop…

Cited by 0SourceScholar
2018

Estimation of the Sound Field at Arbitrary Positions in Distributed Microphone Networks Based on Distributed Ray Space Transform

ICASSP 2018accepted

In this paper we propose a parametric sound field reconstruction approach. In particular, the technique is based on the estimation of three parameters for each acoustic source (source position, radiation pattern and source signal) given the signals acquired by few arbitrarily placed microphone array…

Cited by 0SourceScholar
2017

Dictionary-based Equivalent Source Method for Near-Field Acoustic Holography

ICASSP 2017accepted

In this paper, we propose a modification of the standard Equivalent Source Method (ESM) for Near-Field Acoustic Holography (NAH). As in EMS, we aim at modeling the acoustic pressure radiated from a vibrating object, and its surface velocity, as the joint effect of a set of equivalent sources located…

Cited by 0SourceScholar
2016

A linear operator for the computation of soundfield maps

ICASSP 2016accepted

In the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be convenie…

Cited by 0SourceScholar
2016

A low-cost solution to 3D pinna modeling for HRTF prediction

ICASSP 2016accepted

We propose an infrared (IR) stereo-vision system for estimating the 3D model of the pinna, based on low-cost devices. A commercial IR calibrated stereo camera is used in conjunction with a structured IR light projector, to acquire highly textured snapshots of the pinna. A point cloud is computed for…

Cited by 0SourceScholar