← Search

Erik Marchi

14 accepted papers

2025

SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions

ICASSP 2025accepted

In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two auxiliary tasks related to interactions with virtual assistants…

Cited by 0SourceScholar
2024

A Multimodal Approach to Device-Directed Speech Detection with Large Language Models

ICASSP 2024accepted

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore t…

Cited by 0SourceScholar
2023

Audio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASR

ICASSP 2023accepted

Accurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. a smartphone) is critical for achieving naturalistic, engaging, and privacy-centric interactions with the VA. To this end, we present a novel approach to predict the user intention (whether the user is s…

Cited by 0SourceScholar
2023

Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types

ICASSP 2023accepted

Suppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to…

Cited by 0SourceScholar
2021

Knowledge Transfer for Efficient on-Device False Trigger Mitigation

ICASSP 2021accepted

In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as a "false trigger" and false trigger mitigation (FTM) is essential for designing a privacy-centric non-intrusive smart as…

Cited by 0SourceScholar
2021

On The Role of Visual Cues in Audiovisual Speech Enhancement

ICASSP 2021accepted

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual cues provide not only high-level information abou…

Cited by 0SourceScholar
2021

Progressive Voice Trigger Detection: Accuracy vs Latency

ICASSP 2021accepted

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by including more audio context after a detected trigger phrase, we can indeed get a more…

Cited by 0SourceScholar
2020

Detecting Emotion Primitives from Speech and Their Use in Discerning Categorical Emotions

ICASSP 2020accepted

Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely heavily on speech recognition and natural language understanding for speech content understanding, the investigation of…

Cited by 0SourceScholar
2020

Generating Multilingual Voices Using Speaker Space Translation Based on Bilingual Speaker Data

ICASSP 2020accepted

We present progress towards bilingual Text-to-Speech which is able to transform a monolingual voice to speak a second language while preserving speaker voice quality. We demonstrate that a bilingual speaker embedding space contains a separate distribution for each language and that a simple transfor…

Cited by 0SourceScholar
2020

Multi-Task Learning for Speaker Verification and Voice Trigger Detection

ICASSP 2020accepted

Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single network to perform both tasks jointly. We train the network in a supervised multi-task learning setup, where the speech tr…

Cited by 0SourceScholar
2018

Generalised Discriminative Transform via Curriculum Learning for Speaker Recognition

ICASSP 2018accepted

In this paper we introduce a speaker verification system deployed on mobile devices that can be used to personalise a keyword spotter. We describe a baseline DNN system that maps an utterance to a speaker embedding, which is used to measure speaker differences via cosine similarity. We then introduc…

Cited by 0SourceScholar
2016

Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network

ICASSP 2016accepted

The automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers whi…

Cited by 0SourceScholar
2016

Enhanced semi-supervised learning for multimodal emotion recognition

ICASSP 2016accepted

Semi-Supervised Learning (SSL) techniques have found many applications where labeled data is scarce and/or expensive to obtain. However, SSL suffers from various inherent limitations that limit its performance in practical applications. A central problem is that the low performance that a classifier…

Cited by 0SourceScholar
2015

A novel approach for automatic acoustic novelty detection using a denoising autoencoder with bidirectional LSTM neural networks

ICASSP 2015accepted

Acoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel unsupervised approach based on a denoising autoencoder. In our approach auditory spectral features are process…

Cited by 0SourceScholar