← Search

Yang Ai

13 accepted papers

2026

Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding

AAAI 2026technical

Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over tim

Cited by 0SourcePDFScholar
2025

A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification

ICASSP 2025accepted

In this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local…

Cited by 0SourceScholar
2025

Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations

ICASSP 2025accepted

In this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features…

Cited by 0SourceScholar
2025

CASC-XVC: Zero-Shot Cross-Lingual Voice Conversion with Content Accordant and Speaker Contrastive Losses

ICASSP 2025accepted

Cross-lingual voice conversion (XVC) is a technology that modifies speaker identity while preserving linguistic content in scenarios where the source and target speakers use different languages. Previous non-parallel disentanglement-based methods face severe training-testing inconsistency issues in…

Cited by 0SourceScholar
2025

Can Automated Speech Recognition Errors Provide Valuable Clues for Alzheimer's Disease Detection?

ICASSP 2025accepted

Recent advances in automatic speech recognition (ASR) technology have boosted the viability of fully automated Alzheimer’s disease (AD) detection via ASR transcripts. However, there is a lack of understanding of how ASR errors affect the performance of AD detection. This paper addresses that gap. Fi…

Cited by 0SourceScholar
2025

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis

ICASSP 2025accepted

This paper proposes an Incremental Disentanglement-based Environment-Aware zero-shot text-to-speech (TTS) method, dubbed IDEA-TTS, that can synthesize speech for unseen speakers while preserving the acoustic characteristics of a given environment reference speech. IDEA-TTS adopts VITS as the TTS bac…

Cited by 0SourceScholar
2025

Recursive Feature Learning from Pre-Trained Models for Spoofing Speech Detection

ICASSP 2025accepted

It was recently revealed that using features extracted from pre-trained models can achieve much better performance than using conventional hand-crafted acoustic features for spoofing speech detection. In this paper, we therefore enhance the features from pre-trained model based on recursive learning…

Cited by 0SourceScholar
2024

Considering Temporal Connection between Turns for Conversational Speech Synthesis

ICASSP 2024accepted

Conversational speech synthesis aims to synthesize speech of an individual speaker based on history conversation. However, most studies in conversational speech synthesis only focus on the synthesis performance of the current speaker’s turn and neglect the temporal relationship between turns of inte…

Cited by 0SourceScholar
2023

Neural Speech Phase Prediction Based on Parallel Estimation Architecture and Anti-Wrapping Losses

ICASSP 2023accepted

This paper presents a novel speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra by neural networks. The proposed model is a cascade of a residual convolutional network and a parallel estimation architecture. The parallel estimation architecture is compo…

Cited by 0SourceScholar
2023

Speech Reconstruction from Silent Tongue and Lip Articulation by Pseudo Target Generation and Domain Adversarial Training

ICASSP 2023accepted

This paper studies the task of speech reconstruction from ultrasound tongue images and optical lip videos recorded in a silent speaking mode, where people only activate their intra-oral and extra-oral articulators without producing sound. This task falls under the umbrella of articulatory-to-acousti…

Cited by 0SourceScholar
2023

Zero-Shot Personalized Lip-To-Speech Synthesis with Face Image Based Voice Control

ICASSP 2023accepted

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies can not achieve voice control under zero-shot condition, be…

Cited by 0SourceScholar
2019

Dnn-based Spectral Enhancement for Neural Waveform Generators with Low-bit Quantization

ICASSP 2019accepted

This paper presents a spectral enhancement method to improve the quality of speech reconstructed by neural waveform generators with low-bit quantization. At training stage, this method builds a multiple-target DNN, which predicts log amplitude spectra of natural high-bit waveforms together with the…

Cited by 0SourceScholar