← Search

Pengyuan Zhang

21 accepted papers

2025

Debiased Training For Semi-supervised Sound Event Detection

ICASSP 2025accepted

Recently, semi-supervised sound event detection has attracted increasing research interest due to the scarcity of labeled data. However, traditional semi-supervised learning methods can lead to training instability and confirmation bias because of potentially incorrect pseudo labels. To address this…

Cited by 0SourceScholar
2025

SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

ICASSP 2025accepted

Recently, "textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, the generated speech samples often lack semantic coherence. In this paper, we propose SLM and LLM Integration for spontaneo…

Cited by 5SourceScholar
2024

Improving Short Utterance Anti-Spoofing with Aasist2

ICASSP 2024accepted

The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during…

Cited by 0SourceScholar
2024

One-Class Knowledge Distillation for Spoofing Speech Detection

ICASSP 2024accepted

The detection of spoofing speech generated by unseen algorithms remains an unresolved challenge. One reason for the lack of generalization ability is that traditional detecting systems follow the binary classification paradigm, which inherently assumes the possession of prior knowledge of spoofing s…

Cited by 0SourceScholar
2024

One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection Model

ICASSP 2024accepted

The outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing…

Cited by 0SourceScholar
2024

Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs Detection

ICASSP 2024accepted

Obstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of p…

Cited by 0SourceScholar
2023

Multi-Dimensional Frequency Dynamic Convolution with Confident Mean Teacher for Sound Event Detection

ICASSP 2023accepted

Recently, convolutional neural networks (CNNs) have been widely used in sound event detection (SED). However, traditional convolution is deficient in learning time-frequency domain representation of different sound events. To address this issue, we propose multi-dimensional frequency dynamic convolu…

Cited by 0SourceScholar
2023

PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker Verification

ICASSP 2023accepted

ECAPA-TDNN is currently the most popular TDNN-series model for speaker verification, which refreshed the state-of-the-art (SOTA) performance of TDNN models. However, one-dimensional convolution has a global receptive field over the feature channel. It destroys the time-frequency relevance of the spe…

Cited by 0SourceScholar
2023

Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 Detection

ICASSP 2023accepted

A fast and efficient COVID-19 detection method is of vital importance to control the spread of the epidemic. Many studies have achieved good performance on cough-based COVID19 detection in the past two years. However, the effect of position information in time-frequency features of cough audio has b…

Cited by 0SourceScholar
2022

DPT-FSNet: Dual-Path Transformer Based Full-Band and Sub-Band Fusion Network for Speech Enhancement

ICASSP 2022accepted

Sub-band models have achieved promising results due to their ability to model local patterns in the spectrogram. Some studies further improve the performance by fusing sub-band and full-band information. However, the structure for the full-band and sub-band fusion model was not fully explored. This…

Cited by 0SourceScholar
2022

Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models

ICASSP 2022accepted

Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based e…

Cited by 35SourceScholar
2022

Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models

ICASSP 2022accepted

While Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up the decoding process. For real-world deployment, ASR systems are desired to be highly accurate while achieving fast infe…

Cited by 0SourceScholar
2021

History Utterance Embedding Transformer LM for Speech Recognition

ICASSP 2021accepted

History utterances contain rich contextual information; however, better extracting information from the history utterances and using it to improve the language model (LM) is still challenging. In this paper, we propose the history utterance embedding Transformer LM (HTLM), which includes an embeddin…

Cited by 0SourceScholar
2021

Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data

ICASSP 2021accepted

This paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encode…

Cited by 0SourceScholar
2021

RNN-T Based Open-Vocabulary Keyword Spotting in Mandarin with Multi-Level Detection

ICASSP 2021accepted

Despite the recent prevalence of keyword spotting (KWS) in smart-home, open-vocabulary KWS remains a keen but unmet need among the users. In this paper, we propose an RNN Transducer (RNN-T) based keyword spotting system with a constrained attention mechanism biasing module that biases the RNN-T mode…

Cited by 0SourceScholar
2021

The Thinkit System for Icassp2021 M2voc Challenge

ICASSP 2021accepted

In this paper, we introduce the low resource text-to-speech system from the ThinkIT team submitted to Multi-Speaker Multi-Style Voice Cloning Challenge (M2VoC). The challenge has two tasks: few-shot track1 provides 100 samples for each person and one-shot track2 offers 5 samples only. Each track con…

Cited by 0SourceScholar
2020

CN-Celeb: A Challenging Chinese Speaker Recognition Dataset

ICASSP 2020accepted

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limit…

Cited by 271SourceScholar
2020

Transformer-Based Online CTC/Attention End-To-End Speech Recognition Architecture

ICASSP 2020accepted

Recently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contai…

Cited by 0SourceScholar
2019

An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal Module

ICASSP 2019accepted

Deep convolutional neural network (DCNN) has recently improved the performance of acoustic scene classification. However, the input features of the network are usually based on predefined hand-tailored filters, which may not apply to the specific tasks. To overcome this, we propose a hybrid framewor…

Cited by 0SourceScholar
2019

Self-attention Based Prosodic Boundary Prediction for Chinese Speech Synthesis

ICASSP 2019accepted

Predicting prosodic boundaries from input text plays an important role in Chinese text-to-speech (TTS) system, which directly influences the naturalness and intelligibility of synthesized speech. In this paper, we propose to combine self-attention with multitask learning for prosodic boundary predic…

Cited by 0SourceScholar
2018

Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask Learning

ICASSP 2018accepted

Acoustic signals from microphone arrays are used to improve performance in distant speech recognition due to the availability of spatial information. And multichannel automatic speech recognition (ASR) systems often separate speech enhancement module from acoustic modeling, which may be not optimal…

Cited by 0SourceScholar