← Search

Maurizio Omologo

9 accepted papers

2022

A Neural Prosody Encoder for End-to-End Dialogue Act Classification

ICASSP 2022accepted

Dialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful for DAC. Despite their importance, little research has explored neural approaches to integrate prosodic features into en…

Cited by 0SourceScholar
2022

Caching Networks: Capitalizing on Common Speech for ASR

ICASSP 2022accepted

We introduce Caching Networks (CachingNets), a speech recognition network architecture capable of delivering faster, more accurate decoding by leveraging common speech patterns. By explicitly incorporating select sentences unique to each user into the network’s design, we show how to train the model…

Cited by 0SourceScholar
2019

Accurate Target Annotation in 3D from Multimodal Streams

ICASSP 2019accepted

Accurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverag…

Cited by 0SourceScholar
2018

3D Mouth Tracking from a Compact Microphone Array Co-Located with a camera

ICASSP 2018accepted

We address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood c…

Cited by 0SourceScholar
2017

3D audio-visual speaker tracking with an adaptive particle filter

ICASSP 2017accepted

We propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter frame…

Cited by 0SourceScholar
2017

A network of deep neural networks for Distant Speech Recognition

ICASSP 2017accepted

Despite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationary noises and reverberation are met. A prominent limitation of current systems li…

Cited by 0SourceScholar
2015

A multi-channel corpus for distant-speech interaction in presence of known interferences

ICASSP 2015accepted

This paper describes a new corpus of multi-channel audio data designed to study and develop distant-speech recognition systems able to cope with known interfering sounds propagating in an environment. The corpus consists of both real and simulated signals and of a corresponding detailed annotation.…

Cited by 0SourceScholar
2015

Audio source separation using a redundant library of source spectral bases for non-negative tensor factorization

ICASSP 2015accepted

This work proposes a solution to the problem of under-determined audio source separation using pre-trained redundant source-based prior information. In local Gaussian modeling of a mixing process, an observed mixture is modeled by a Gaussian distribution parameterized by source variances and spatial…

Cited by 0SourceScholar