← Search

Pengcheng Guo

17 accepted papers

2025

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

ICASSP 2025accepted

Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for humans and machines to detect. In this study, we p…

Cited by 0SourceScholar
2024

Automatic Channel Selection and Spatial Feature Integration for Multi-Channel Speech Recognition Across Various Array Topologies

ICASSP 2024accepted

Automatic Speech Recognition (ASR) has shown remarkable progress, yet it still faces challenges in real-world distant scenarios across various array topologies each with multiple recording devices. The focal point of the CHiME-7 Distant ASR task is to devise a unified system capable of generalizing…

Cited by 0SourceScholar
2024

Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study

ICASSP 2024accepted

Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features such as spectrograms are often used as the input for the subsequent model. However, they can still be redundant. Recent…

Cited by 0SourceScholar
2024

MLCA-AVSR: Multi-Layer Cross Attention Fusion Based Audio-Visual Speech Recognition

ICASSP 2024accepted

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve the system’s robustness. However, current studies mainly focus on fusing the we…

Cited by 0SourceScholar
2023

Distinguishable Speaker Anonymization Based on Formant and Fundamental Frequency Scaling

ICASSP 2023accepted

Speech data on the Internet are proliferating exponentially because of the emergence of social media, and the sharing of such personal data raises obvious security and privacy concerns. One solution to mitigate these concerns involves concealing speaker identities before sharing speech data, also re…

Cited by 0SourceScholar
2023

NC-WAMKD: Neighborhood Correction Weight-Adaptive Multi-Teacher Knowledge Distillation for Graph-Based Semi-Supervised Node Classification

ICASSP 2023accepted

Multi-teacher knowledge distillation can improve the performance of student networks in semi-supervised node classification tasks, but existing works ignore the importance of different teachers, using average of multiple teachers as final prediction. In addition, they rely on a large amount of label…

Cited by 0SourceScholar
2023

Preserving Background Sound in Noise-Robust Voice Conversion Via Multi-Task Learning

ICASSP 2023accepted

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare attention to VC with background sound. The critical problem for pre…

Cited by 0SourceScholar
2023

The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge

ICASSP 2023accepted

This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. Specifically, the weighted prediction error (WPE) and guided source separation (GSS) techniques are used to reduce rever…

Cited by 0SourceScholar
2023

VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting

ICASSP 2023accepted

The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting, which leverages complementary relationships over multiple moda…

Cited by 0SourceScholar
2022

M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologi…

Cited by 0SourceScholar
2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2022

WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition

ICASSP 2022accepted

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of…

Cited by 0SourceScholar
2021

Recent Developments on Espnet Toolkit Boosted By Conformer

ICASSP 2021accepted

In this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, suc…

Cited by 0SourceScholar
2020

A Multi-Scaled Receptive Field Learning Approach for Medical Image Segmentation

ICASSP 2020accepted

Biomedical image segmentation has been widely studied, and lots of methods have been proposed. Among these methods, attention U-Net has achieved a promising performance. However, it has drawbacks of extracting the multi-scaled receptive field features at the high-level feature maps, resulting in the…

Cited by 0SourceScholar
2020

Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals

NeurIPS 2020poster

Neural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence. In this work, we focus on one-to-many sequence transduction problems, such as extracting multiple sequential sources from a mixture sequence.…

2019

Domain Adversarial Training for Improving Keyword Spotting Performance of ESL Speech

ICASSP 2019accepted

A second language (L2) learner usually cannot speak L2 well in both pronunciations and forming-of-words. Hence his/her L2 speech cannot be well recognized by a recognizer trained with native data. Domain adversarial training (DAT), capable of reducing the acoustic mismatch between training and testi…

Cited by 0SourceScholar