← Search

John H. L. Hansen

33 accepted papers

2025

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

ICASSP 2025accepted

Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for humans and machines to detect. In this study, we p…

Cited by 0SourceScholar
2025

Semi-Supervised Speaker Diarization Using Graph Transformers and LLMs on Naturalistic Apollo 11 Data

ICASSP 2025accepted

Speaker diarization is the process of segmenting and tagging audio streams based on speaker identity. Traditional methods face significant challenges in real-world scenarios when applied to spontaneous multi-speaker conversational speech. The Fearless Steps Apollo 11 corpus (FS-A11) presents real-wo…

Cited by 0SourceScholar
2024

Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection

ICASSP 2024accepted

Speaker diarization has traditionally been explored using datasets that are either clean, feature a limited number of speakers, or have a large volume of data but lack the complexities of real-world scenarios. This study takes a unique approach by focusing on the Fearless Steps APOLLO audio resource…

Cited by 0SourceScholar
2024

Dual-Path Minimum-Phase and All-Pass Decomposition Network for Single Channel Speech Dereverberation

ICASSP 2024accepted

With the development of deep neural networks (DNN), many DNN-based speech dereverberation approaches have been proposed to achieve significant improvement over the traditional methods. However, most deep learning-based dereverberation methods solely focus on suppressing time-frequency domain reverbe…

Cited by 0SourceScholar
2024

Efficient Adapter Tuning of Pre-Trained Speech Models for Automatic Speaker Verification

ICASSP 2024accepted

With excellent generalization ability, self-supervised speech models have shown impressive performance on various downstream speech tasks in the pre-training and fine-tuning paradigm. However, as the growing size of pre-trained models, fine-tuning becomes practically unfeasible due to heavy computat…

Cited by 0SourceScholar
2024

Fearless Steps Apollo: Team Communications Based Community Resource Development for Science, Technology, Education, and Historical Preservation

ICASSP 2024accepted

The Fearless Steps Apollo (FS-APOLLO) resource is a collection of 150,000 hours of audio, associated meta-data, and supplemental speech technology infrastructure intended to benefit the (i) speech processing technology, (ii) communication science, team-based psychology, and (iii) education/STEM, his…

Cited by 0SourceScholar
2024

Situational Signal Processing with Ecological Momentary Assessment: Leveraging Environmental Context for Cochlear Implant Users

ICASSP 2024accepted

Technological advancements for biomedical interfaces and devices, such as cochlear implants (CIs), depend on the integration of novel signal processing strategies enhanced by situational real-time feedback in naturalistic spontaneous environments. This study proposes the first CI framework for situa…

Cited by 0SourceScholar
2024

T-EnFP: An Efficient Transformer Encoder-Based System for Driving Behavior Classification

ICASSP 2024accepted

Recently, Transformer-based architectures have been explored for classifying driving behavior. Although the Transformer effectively employs self-attention for global temporal learning, the presence of redundant modules can detrimentally affect task-specific performance and overall efficiency. In thi…

Cited by 0SourceScholar
2023

Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting

ICASSP 2023accepted

In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels i…

Cited by 0SourceScholar
2023

Improving Transformer-Based Networks with Locality for Automatic Speaker Verification

ICASSP 2023accepted

Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embeddings, it is inadequate for capturing short-range local context, which is essent…

Cited by 0SourceScholar
2021

DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning

ICASSP 2021accepted

Despite speaker verification has achieved significant performance improvement with the development of deep neural networks, do-main mismatch is still a challenging problem in this field. In this study, we propose a novel framework to disentangle speaker-related and domain-specific features and apply…

Cited by 0SourceScholar
2020

A multi-view approach for Mandarin non-native mispronunciation verification

ICASSP 2020accepted

Traditionally, the performance of non-native mispronunciation verification systems relied on effective phone-level labelling of non-native corpora. In this study, a multi-view approach is proposed to incorporate discriminative feature representations which requires less annotation for non-native mis…

Cited by 0SourceScholar
2019

Cross-lingual Text-independent Speaker Verification Using Unsupervised Adversarial Discriminative Domain Adaptation

ICASSP 2019accepted

Speaker verification systems often degrade significantly when there is a language mismatch between training and testing data. Being able to improve cross-lingual speaker verification system using unlabeled data can greatly increase the robustness of the system and reduce human labeling costs. In thi…

Cited by 0SourceScholar
2019

Semi-supervised Learning with Generative Adversarial Networks for Arabic Dialect Identification

ICASSP 2019accepted

Dialect Identification (DID) refers to the process of identifying different dialects within the same language class. Compared with more general language identification (LID), DID is a more challenging task because of the substantial similarity between dialects. For an i-vector based LID/DID, prior s…

Cited by 0SourceScholar
2019

Transfer Learning Using Raw Waveform Sincnet for Robust Speaker Diarization

ICASSP 2019accepted

Speaker diarization tells who spoke and when? in an audio stream. SincNet is a recently developed novel convolutional neural network (CNN) architecture where the first layer consists of parameterized sinc filters. Unlike conventional CNNs, SincNet take raw speech waveform as input. This paper levera…

Cited by 0SourceScholar
2019

UTD-CRSS Systems for 2018 NIST Speaker Recognition Evaluation

ICASSP 2019accepted

In this study, we present systems submitted by the Center for Robust Speech Systems (CRSS) from UTDallas to NIST SRE 2018 (SRE18). Three alternative front-end speaker embedding frameworks are investigated, that includes: (i) i-vector, (ii) x-vector, (iii) and a modified triplet speaker embedding sys…

Cited by 0SourceScholar
2018

Robust Feature Clustering for Unsupervised Speech Activity Detection

ICASSP 2018accepted

In certain applications such as zero-resource speech processing or very-low resource speech-language systems, it might not be feasible to collect speech activity detection (SAD) annotations. However, the state-of-the-art supervised SAD techniques based on neural networks or other machine learning me…

Cited by 0SourceScholar
2017

A study of speaker verification performance with expressive speech

ICASSP 2017accepted

Expressive speech introduces variations in the acoustic features affecting the performance of speech technology such as speaker verification systems. It is important to identify the range of emotions for which we can reliably estimate speaker verification tasks. This paper studies the performance of…

Cited by 0SourceScholar
2017

Environment aware speaker diarization for moving targets using parallel DNN-based recognizers

ICASSP 2017accepted

Current diarization algorithms are commonly applied to the outputs of single non-moving microphones. They do not explicitly identify the content of overlapped segments from multiple speakers or acoustic events. This paper presents an acoustic environment aware child-adult diarization applied to the…

Cited by 0SourceScholar
2017

i-Vector/PLDA speaker recognition using support vectors with discriminant analysis

ICASSP 2017accepted

i-Vector feature representation with probabilistic linear discriminant analysis (PLDA) scoring in speaker recognition system has recently achieved effective permanence even on channel mismatch conditions. In general, experiments carried out using this combined strategy employ linear discriminant ana…

Cited by 0SourceScholar
2016

F0 estimation for noisy speech by exploring temporal harmonic structures in local time frequency spectrum segment

ICASSP 2016accepted

In this paper, we propose a noise robust F0 estimation approach by exploring the temporal harmonic structures in local time-frequency (TF) spectrum segment. Since the speech energy is sparsely distributed on the TF plane, the speech harmonic structures occupied in the higher speech energy TF segment…

Cited by 0SourceScholar
2016

Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach

ICASSP 2016accepted

Sustaining automatic speaker verification(ASV) systems from spoofing attacks remains an essential challenge, even if significant progress in ASV has been achieved in recent years. In this study, an automatic spoofing detection approach using an i-vector framework is proposed. Two approaches are used…

Cited by 0SourceScholar
2016

Language recognition using deep neural networks with very limited training data

ICASSP 2016accepted

This study proposes a novel deep neural network (DNN) based approach to language identification (LID) for the NIST 2015 Language Recognition (LRE) i-Vector Machine Learning Challenge. State-of-the-art DNN based LID systems utilize large amounts of labeled training data. The 2015 LRE i-Vector Machine…

Cited by 0SourceScholar
2016

UTD-CRSS system for the NIST 2015 language recognition i-vector machine learning challenge

ICASSP 2016accepted

In this paper, we present the system developed by the Center for Robust Speech Systems (CRSS), University of Texas at Dallas, for the NIST 2015 language recognition i-vector machine learning challenge. Our system includes several subsystems, based on Linear Discriminant Analysis - Support Vector Mac…

Cited by 0SourceScholar
2015

Analysis of speech and language communication for cochlear implant users in noisy lombard conditions

ICASSP 2015accepted

Acoustic/linguistic modification of speech production with respect to auditory feedback is an important research domain for robust human-to-human and human-to-machine communication. For instance, in the presence of environmental noise, a speaker experiences the well-known phenomenon termed as Lombar…

Cited by 0SourceScholar
2015

Generative modeling of pseudo-target domain adaptation samples for whispered speech recognition

ICASSP 2015accepted

The lack of available large corpora of transcribed whispered speech is one of the major roadblocks for development of successful whisper recognition engines. Our recent study has introduced a Vector Taylor Series (VTS) approach to pseudo-whisper sample generation which requires availability of only…

Cited by 0SourceScholar
2015

Image-guided customization of frequency-place mapping in cochlear implants

ICASSP 2015accepted

Multi-channel cochlear implants (CI) leverage frequency based cochlear tonotopic mapping to map acoustic information to the cochlear place of stimulation which is primarily determined by electrode locations. Despite the fact that electrode locations within the cochlea are unique to each patient, the…

Cited by 0SourceScholar
2015

Leveraging automatic speech recognition in cochlear implants for improved speech intelligibility under reverberation

ICASSP 2015accepted

Despite recent advancements in digital signal processing technology for cochlear implant (CI) devices, there still remains a significant gap between speech identification performance of CI users in reverberation compared to that in anechoic quiet conditions. Alternatively, automatic speech recogniti…

Cited by 0SourceScholar
2015

Prof-Life-Log: Analysis and classification of activities in daily audio streams

ICASSP 2015accepted

A new method to analyze and classify daily activities in personal audio recordings (PARs) is presented. The method employs speech activity detection (SAD) and speaker diarization systems to provide high level semantic segmentation of the audio file. Subsequently, a number of audio, speech and lexica…

Cited by 0SourceScholar
2015

Robust overlapped speech detection and its application in word-count estimation for Prof-Life-Log data

ICASSP 2015accepted

The ability to estimate the number of words spoken by an individual over a certain period of time is valuable in second language acquisition, healthcare, and assessing language development. However, establishing a robust automatic framework to achieve high accuracy is non-trivial in realistic/natura…

Cited by 0SourceScholar
2015

Robust unsupervised detection of human screams in noisy acoustic environments

ICASSP 2015accepted

This study is focused on an unsupervised approach for detection of human scream vocalizations from continuous recordings in noisy acoustic environments. The proposed detection solution is based on compound segmentation, which employs weighted mean distance, T <sup xmlns:mml="http://www.w3.org/1998/M…

Cited by 0SourceScholar
2015

Weighted training for speech under Lombard Effect for speaker recognition

ICASSP 2015accepted

The presence of Lombard Effect in speech is proven to have severe effects on the performance of speech systems, especially speaker recognition. Varying kinds of Lombard speech are produced by speakers under influence of varying noise types [1]. This study proposes a high-accuracy classifier using de…

Cited by 0SourceScholar