← Search

Thomas Fang Zheng

8 accepted papers

2025

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

ICASSP 2025accepted

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowdsourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours…

Cited by 2SourceScholar
2024

Enhancing Quantised End-to-End ASR Models Via Personalisation

ICASSP 2024accepted

Recent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Model quantisation is an effective solution that sometimes causes the word error rate (WER) to increase. In this paper, a…

Cited by 4SourceScholar
2023

CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis

ICASSP 2023accepted

Research on Video to Speech Synthesis (VTS) surges recently and the focus is gradually shifting from small-vocabulary short-phrase VTS to large-vocabulary continuous VTS (LVC-VTS). A large-scale dataset with sufficient speakers and utterances is a prerequisite for such research, and the database is…

Cited by 0SourceScholar
2021

Attack on Practical Speaker Verification System Using Universal Adversarial Perturbations

ICASSP 2021accepted

In authentication scenarios, applications of practical speaker verification systems usually require a person to read a dynamic authentication text. Previous studies played an audio adversarial example as a digital signal to perform physical attacks, which would be easily rejected by audio replay det…

Cited by 0SourceScholar
2021

Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification

ICASSP 2021accepted

Domain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly p…

Cited by 0SourceScholar
2017

Speaker segmentation using deep speaker vectors for fast speaker change scenarios

ICASSP 2017accepted

A novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a kind of frame-level speaker discriminative feature, whose dis…

Cited by 0SourceScholar