← Search

Tiago H. Falk

12 accepted papers

2024

VIC-KD: Variance-Invariance-Covariance Knowledge Distillation to Make Keyword Spotting More Robust Against Adversarial Attacks

ICASSP 2024accepted

Keyword spotting (KWS) refers to the task of identifying a set of predefined words in audio streams. With the advances seen recently with deep neural networks, it has become a popular technology to activate and control small devices, such as voice assistants. Relying on such models for edge devices,…

Cited by 0SourceScholar
2023

Robustdistiller: Compressing Universal Speech Representations for Enhanced Environment Robustness

ICASSP 2023accepted

Self-supervised speech pre-training enables deep neural network models to capture meaningful and disentangled factors from raw waveform signals. The learned universal speech representations can then be used across numerous down-stream tasks. These representations, however, are sensitive to distribut…

Cited by 15SourceScholar
2022

Fusion of Modulation Spectral and Spectral Features with Symptom Metadata for Improved Speech-Based Covid-19 Detection

ICASSP 2022accepted

Existing speech-based coronavirus disease 2019 (COVID-19) detection systems provide poor interpretability and limited robustness to unseen data conditions. In this paper, we propose a system to overcome these limitations. In particular, we propose to fuse two different feature modalities with patien…

Cited by 11SourceScholar
2021

Context-Aware Speech Stress Detection in Hospital Workers Using Bi-LSTM Classifiers

ICASSP 2021accepted

Hospital workers are known to work long hours in a highly stressful environment. The COVID-19 pandemic has increased this burden multi-fold. Pre-COVID statistics already showed that one in every three nurses reported burnout, thus affecting patient satisfaction and the quality of their provided serv…

Cited by 0SourceScholar
2020

An Ensemble Based Approach for Generalized Detection of Spoofing Attacks to Automatic Speaker Recognizers

ICASSP 2020accepted

As automatic speaker recognizer systems become mainstream, voice spoofing attacks are on the rise. Common attack strategies include replay, the use of text-to-speech synthesis, and voice conversion systems. While previouslyproposed end-to-end detection frameworks have shown to be effective in spotti…

Cited by 0SourceScholar
2018

Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio Estimation

ICASSP 2018accepted

Non-intrusive estimators for acoustic parameters like the direct-to-reverberation ratio (DRR) are useful tools but still perform weakly as shown in the acoustic characterization of environments (ACE) challenge. In this paper, we develop a novel dual-channel metric based on the modulation energy doma…

Cited by 0SourceScholar
2018

Improved Audio-Visual Laughter Detection Via Multi-Scale Multi-Resolution Image Texture Features and Classifier Fusion

ICASSP 2018accepted

Efforts are afoot to design better context-aware human-computer interaction techniques that have knowledge of both their surrounding and the affective state of the user. One of the most important nonverbal behavioural cues for affective human-machine interaction is laughter. Automatic detection of l…

Cited by 0SourceScholar
2017

Event-related synchronisation responses to N-back memory tasks discriminate between healthy ageing, mild cognitive impairment, and mild Alzheimer's disease

ICASSP 2017accepted

In this study we investigate whether or not event-related (de)synchronisation (ERD/ERS) can be used to differentiate between 27 healthy elderly, 21 subjects diagnosed with amnestic mild cognitive impairment (aMCI) and 16 mild Alzheimer's disease (AD) patients. Using 32-channel EEG recordings, we mea…

Cited by 0SourceScholar
2017

Speech temporal dynamics fusion approaches for noise-robust reverberation time estimation

ICASSP 2017accepted

Reverberation and noise are known to be the two most important culprits for poor performance in far-field speech applications, such as automatic speech recognition. Recent research has suggested that reverberation-aware speech enhancement (or speech technologies, in general) could be used to improve…

Cited by 5SourceScholar
2016

Feature mapping, score-, and feature-level fusion for improved normal and whispered speech speaker verification

ICASSP 2016accepted

In this paper, automatic speaker verification using normal and whispered speech is explored. Typically, for speaker verification systems with varying vocal effort inputs, standard solutions such as feature mapping or addition of data during parameter estimation (training) and enrollment stages resul…

Cited by 0SourceScholar
2015

On the potential for artificial bandwidth extension of bone and tissue conducted speech: a mutual information study

ICASSP 2015accepted

To enhance the communication experience of workers equipped with hearing protection devices and radio communication in noisy environments, alternative methods of speech capture have been utilized. One such approach uses speech captured by a microphone in an occluded ear canal. Although high in signa…

Cited by 0SourceScholar