← Search

Danilo Comminiello

18 accepted papers

2026

Closing the Modality Gap Aligns Group-Wise Semantics

ICLR 2026poster

In multimodal learning, CLIP has been recognized as the \textit{de facto} method for learning a shared latent space across multiple modalities, placing similar representations close to each other and moving them away from dissimilar ones. Although CLIP-based losses effectively align modalities at th…

Cited by 0SourcecodeScholar
2026

TRAINING-FREE MULTIMODAL GUIDANCE FOR VIDEO TO AUDIO GENERATION

ICASSP 2026poster

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results, existing approaches either require costly joint training on larg…

Cited by 0SourcePDFScholar
2025

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity

NeurIPS 2025poster

Multimodal learning plays a pivotal role in advancing artificial intelligence systems by incorporating information from multiple modalities to build a more comprehensive representation. Despite its importance, current state-of-the-art models still suffer from severe limitations that prevent the succ…

Cited by 0SourceScholar
2025

Gramian Multimodal Representation Learning and Alignment

ICLR 2025poster

Human perception integrates multiple modalities—such as vision, hearing, and language—into a unified understanding of the surrounding reality. While recent multimodal models have achieved significant progress by aligning pairs of modalities via contrastive learning, their solutions are unsuitable wh…

2025

Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models

ICASSP 2025accepted

Generating images from brain waves is gaining increasing attention due to its potential to advance brain-computer interface (BCI) systems by understanding how brain signals encode visual cues. Most of the literature has focused on fMRI-to-Image tasks as fMRI is characterized by high spatial resoluti…

Cited by 0SourceScholar
2024

Diffusion Models for Audio Semantic Communication

ICASSP 2024accepted

Directly sending audio signals from a transmitter to a receiver across a noisy channel may absorb consistent bandwidth and be prone to errors when trying to recover the transmitted bits. On the contrary, the recent semantic communication approach proposes to send the semantics and then regenerate se…

Cited by 0SourceScholar
2024

Efficient Functional Link Adaptive Filters Based On Nearest Kronecker Product Decomposition

ICASSP 2024accepted

Functional link adaptive filters (FLAFs) utilize expansion blocks to nonlinearly augment the input signal to a higher dimensional space, after which an adaptive weight algorithm is applied. These filters are useful for nonlinear system identification tasks, as they can update a large number of coeff…

Cited by 1SourceScholar
2024

Enhancing Semantic Communication with Deep Generative Models: An Overview

ICASSP 2024accepted

Semantic communication is poised to play a pivotal role in shaping the landscape of future AI-driven communication systems. Its challenge of extracting semantic information from the original complex content and regenerating semantically consistent data at the receiver, possibly being robust to chann…

Cited by 0SourceScholar
2024

Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley Synthesis

ICASSP 2024accepted

Sound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio with video. In some cases, environmental recordings from video…

Cited by 0SourceScholar
2023

Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality

ICASSP 2023accepted

The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D Sound Event Localization and Detection in Extended Reality ap…

Cited by 0SourceScholar
2022

L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment

ICASSP 2022accepted

The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition <sup xmlns:mml="http://www.w3.org/1998/Math…

Cited by 62SourceScholar
2020

Differentiable Branching In Deep Networks for Fast Inference

ICASSP 2020accepted

In this paper, we consider the design of deep neural networks augmented with multiple auxiliary classifiers departing from the main (backbone) network. These classifiers can be used to perform early-exit from the network at various layers, making them convenient for energy-constrained applications s…

Cited by 0SourceScholar
2019

Frequency-domain Adaptive Filtering: from Real to Hypercomplex Signal Processing

ICASSP 2019accepted

Frequency-domain adaptive filters (FDAFs) have been widely used over the years, but they are still matter of research due to their powerful capabilities that differentiate them from the whole family of time-domain adaptive filters. This paper aims at providing an overview on FDAFs through a unifying…

Cited by 0SourceScholar
2019

Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound Events

ICASSP 2019accepted

Learning from data in the quaternion domain enables us to exploit internal dependencies of 4D signals and treating them as a single entity. One of the models that perfectly suits with quaternion-valued data processing is represented by 3D acoustic signals in their spherical harmonics decomposition.…

Cited by 0SourceScholar
2019

Widely Linear Kernels for Complex-valued Kernel Activation Functions

ICASSP 2019accepted

Complex-valued neural networks (CVNNs) have been shown to be powerful nonlinear approximators when the input data can be properly modeled in the complex domain. One of the major challenges in scaling up CVNNs in practice is the design of complex activation functions. Recently, we proposed a novel fr…

Cited by 0SourceScholar