← Search

Jiangyan Yi

28 accepted papers

2026

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

ICASSP 2026oral

Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level…

Cited by 0SourcePDFScholar
2025

Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization

ICASSP 2025accepted

Partially deepfake audio localization is important in audio forensics. However, existing localization models for partially deepfake audio face two major challenges: distribution shifts between training and testing data as well as insufficient utilization of information from both manipulated regions…

Cited by 0SourceScholar
2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

ICML 2025oral

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suff…

2025

Code-switching Mediated Sentence-level Semantic Learning

AAAI 2025technical

Code-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic…

Cited by 0SourcePDFScholar
2025

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

ICML 2025poster

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-apprai…

2025

PET: High-Frequency Temporal Self-Consistency Learning for Partially Deepfake Audio Localization

ICASSP 2025accepted

Partially deepfake audio attacks have attracted the attention recently, and the demand for locating the manipulation regions of partially deepfake audio arises accordingly. However, existing methods are usually proposed based on frame-level authenticity detection or splicing boundaries detection, ne…

Cited by 0SourceScholar
2025

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

AAAI 2025technical

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real…

2025

WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification

ICASSP 2025accepted

Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training pr…

Cited by 0SourceScholar
2024

Fewer-Token Neural Speech Codec with Time-Invariant Codes

ICASSP 2024accepted

Language model based text-to-speech (TTS) models, like VALL-E, have gained attention for their outstanding in-context learning capability in zero-shot scenarios. Neural speech codec is a critical component of these models, which can convert speech into discrete token representations. However, excess…

Cited by 0SourceScholar
2024

Multi-Scale Permutation Entropy for Audio Deepfake Detection

ICASSP 2024accepted

With the widespread application of Automatic Speaker Verification (ASV) technology in security authentication, the threat of fake audio attacks looms as a malicious means compromising system security. In this study, we employ the multi-scale permutation entropy (MPE) in audio deepfake detection, whi…

Cited by 0SourceScholar
2024

NLoPT: N-gram Enhanced Low-Rank Task Adaptive Pre-training for Efficient Language Model Adaption

COLING 2024main

Pre-trained Language Models (PLMs) like BERT have achieved superior performance on different downstream tasks, even when such a model is trained on a general domain. Moreover, recent studies have shown that continued pre-training on task-specific data, known as task adaptive pre-training (TAPT), can…

Cited by 1SourcePDFScholar
2024

What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection

AAAI 2024technical

The rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms. Existing detection models have shown remarkable success in discriminating known de…

Cited by 29SourcePDFScholar
2023

Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection

ICML 2023poster

Current fake audio detection algorithms have achieved promising performances on most datasets. However, their performance may be significantly degraded when dealing with audio of a different dataset. The orthogonal weight modification to overcome catastrophic forgetting does not consider the similar…

2023

GCC-Speaker: Target Speaker Localization with Optimal Speaker-Dependent Weighting in Multi-Speaker Scenarios

ICASSP 2023accepted

Existing noise-robust and reverberant-robust localization algorithms fail to localize the target speaker when interfering speakers are present. In this paper, we address the problem of localizing only the target speaker in multi-speaker scenarios and propose a target speaker localization algorithm,…

Cited by 0SourceScholar
2023

Learning From Yourself: A Self-Distillation Method For Fake Speech Detection

ICASSP 2023accepted

In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on,…

Cited by 0SourceScholar
2022

A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature

ICASSP 2022accepted

There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose…

Cited by 0SourceScholar
2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2022

Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing

ICASSP 2022accepted

The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records…

Cited by 0SourceScholar
2021

Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis

ICASSP 2021accepted

End-to-end framework can generate high-quality and high-similarity speech in the personalized speech synthesis task. However, the generalization of out-of-domain texts is still a challenging task. Limited target data leads to unacceptable errors and poor prosody and similarity performance of the syn…

Cited by 0SourceScholar
2021

Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition

ICASSP 2021accepted

Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text…

Cited by 0SourceScholar
2021

Patnet : A Phoneme-Level Autoregressive Transformer Network for Speech Synthesis

ICASSP 2021accepted

Aiming at efficiently predicting acoustic features with high naturalness and robustness, this paper proposes PATNet, a neural acoustic model for speech synthesis using phoneme-level autoregression. PATNet accepts phoneme sequences as input and is built based on Transformer structure. PATNet adopts a…

Cited by 0SourceScholar
2021

Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021

ICASSP 2021accepted

The paper describes the CASIA speech synthesis system entry for challenge M2VoC 2021. The low similarity and naturalness of synthesized speech remains a challenging problem for speaker adaptation with few resources. Since the end-to-end acoustic model is too complex to interpret, overfitting will oc…

Cited by 0SourceScholar
2020

Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech Synthesis

ICASSP 2020accepted

End-to-end speech synthesis can generate high-quality synthetic speech and achieve high similarity scores with low-resource adaptation data. However, the generalization of out-domain texts is still a challenging task. The limited adaptation data leads to unacceptable errors and the poor prosody perf…

Cited by 0SourceScholar
2020

Synchronous Transformers for end-to-end Speech Recognition

ICASSP 2020accepted

For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition…

Cited by 0SourceScholar
2019

Language-invariant Bottleneck Features from Adversarial End-to-end Acoustic Models for Low Resource Speech Recognition

ICASSP 2019accepted

This paper proposes to learn language-invariant bottleneck features from an adversarial end-to-end acoustic model for low resource languages. The multilingual end-to-end model is trained with a connectionist temporal classification loss function. The model has shared and private layers. The shared l…

Cited by 0SourceScholar
2018

End-to-End Continuous Emotion Recognition from Video Using 3D Convlstm Networks

ICASSP 2018accepted

Conventional continuous emotion recognition consists of feature extraction step followed by regression step. However, the objective of the two steps is not consistent as they are parted. Besides, there is still no consensus about appropriate emotional features. In this study, we propose an end-to-en…

Cited by 0SourceScholar