← Search

Siddharth Sigtia

7 accepted papers

2025

SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions

ICASSP 2025accepted

In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two auxiliary tasks related to interactions with virtual assistants…

Cited by 0SourceScholar
2024

A Multimodal Approach to Device-Directed Speech Detection with Large Language Models

ICASSP 2024accepted

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore t…

Cited by 0SourceScholar
2021

Progressive Voice Trigger Detection: Accuracy vs Latency

ICASSP 2021accepted

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by including more audio context after a detected trigger phrase, we can indeed get a more…

Cited by 0SourceScholar
2020

Multi-Task Learning for Speaker Verification and Voice Trigger Detection

ICASSP 2020accepted

Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single network to perform both tasks jointly. We train the network in a supervised multi-task learning setup, where the speech tr…

Cited by 0SourceScholar
2020

Multi-Task Learning for Voice Trigger Detection

ICASSP 2020accepted

We describe the design of a voice trigger detection system for smart speakers. In this study, we address two major challenges. The first is that the detectors are deployed in complex acoustic environments with external noise and loud playback by the device itself. Secondly, collecting training examp…

Cited by 0SourceScholar
2018

Generalised Discriminative Transform via Curriculum Learning for Speaker Recognition

ICASSP 2018accepted

In this paper we introduce a speaker verification system deployed on mobile devices that can be used to personalise a keyword spotter. We describe a baseline DNN system that maps an utterance to a speaker embedding, which is used to measure speaker differences via cosine similarity. We then introduc…

Cited by 0SourceScholar
2015

A hybrid recurrent neural network for music transcription

ICASSP 2015accepted

We investigate the problem of incorporating higher-level symbolic score-like information into Automatic Music Transcription (AMT) systems to improve their performance. We use recurrent neural networks (RNNs) and their variants as music language models (MLMs) and present a generative architecture for…

Cited by 0SourceScholar