← Search

Wu Guo

19 accepted papers

2025

A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification

ICASSP 2025accepted

In this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local…

Cited by 0SourceScholar
2025

Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations

ICASSP 2025accepted

In this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features…

Cited by 0SourceScholar
2025

Recursive Feature Learning from Pre-Trained Models for Spoofing Speech Detection

ICASSP 2025accepted

It was recently revealed that using features extracted from pre-trained models can achieve much better performance than using conventional hand-crafted acoustic features for spoofing speech detection. In this paper, we therefore enhance the features from pre-trained model based on recursive learning…

Cited by 0SourceScholar
2024

Generating High-Quality Adversarial Examples with Universal Perturbation-Based Adaptive Network and Improved Perceptual Loss

ICASSP 2024accepted

Deep neural network-based speaker identification systems are vulnerable to adversarial attacks. However, the distortions of the adversarial examples are still obvious in most cases. In this work, we therefore propose a universal perturbation-based adaptive network (UPAN) to generate high-quality adv…

Cited by 0SourceScholar
2024

Meta Representation Learning Method for Robust Speaker Verification in Unseen Domains

ICASSP 2024accepted

This paper presents a meta representation learning method for robust speaker verification (SV) in unseen domains. It is known that the existing embedding learning based SV systems may suffer from domain mismatch issues. To address this, we propose an episodic training procedure to compensate domain…

Cited by 0SourceScholar
2024

Robust Spoof Speech Detection Based on Multi-Scale Feature Aggregation and Dynamic Convolution

ICASSP 2024accepted

Spoof speech detection (SSD) can help to protect an automatic speaker recognition system against malicious attacks. However, there exists a great diversity in the spoof utterances generated by different text-to-speech and voice conversion algorithms, resulting in a poor generality of an SSD system t…

Cited by 0SourceScholar
2023

FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization

EMNLP 2023long main

Summaries of medical text shall be faithful by being consistent and factual with source inputs, which is an important but understudied topic for safety and efficiency in healthcare. In this paper, we investigate and improve faithfulness in summarization on a broad range of medical summarization task…

Cited by 0SourcecodeScholar
2023

Pre-training Language Model as a Multi-perspective Course Learner

ACL 2023findings

ELECTRA, the generator-discriminator pre-training framework, has achieved impressive semantic construction capability among various downstream tasks. Despite the convincing performance, ELECTRA still faces the challenges of monotonous training and deficient interaction. Generator with only masked la…

Cited by 1SourcePDFScholar
2022

Wider & Closer: Mixture of Short-channel Distillers for Zero-shot Cross-lingual Named Entity Recognition

EMNLP 2022main

Zero-shot cross-lingual named entity recognition (NER) aims at transferring knowledge from annotated and rich-resource data in source languages to unlabeled and lean-resource data in target languages. Existing mainstream methods based on the teacher-student distillation framework ignore the rich and…

2020

An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales

ICASSP 2020accepted

This paper presents an improved deep embedding learning method based on a convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) a multiscale convolution (MSCNN) is adopted in the frame-level layers to capture…

Cited by 0SourceScholar
2020

Attention-Based Gated Scaling Adaptive Acoustic Model for CTC-Based Speech Recognition

ICASSP 2020accepted

In this paper, we propose a novel adaptive technique that uses an attention-based gated scaling (AGS) scheme to improve deep feature learning for connectionist temporal classification (CTC) acoustic modeling. In AGS, the outputs of each hidden layer of the main network are scaled by an auxiliary gat…

Cited by 0SourceScholar
2019

A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification

ICASSP 2019accepted

Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challen…

Cited by 0SourceScholar
2019

Topic Detection in Conversational Telephone Speech Using CNN with Multi-stream Inputs

ICASSP 2019accepted

Topic detection for conversational telephone speech (CTS) is addressed in this paper. The low accuracy of automatic speech recognition (ASR) will cause severe performance deterioration for topic detection. To make up for this, we adopt two ASR systems, HMM-BiLSTM and CTC systems, to provide compleme…

Cited by 0SourceScholar
2018

Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis

ICASSP 2018accepted

In recent years, neural networks (NN) have achieved remarkable performance improvement in text classification due to their powerful ability to encode discriminative features by incorporating label information into model training. Inspired by the success of NN in text classification, we propose a pse…

Cited by 0SourceScholar
2015

Channel adaptation of plda for text-independent speaker verification

ICASSP 2015accepted

Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log…

Cited by 0SourceScholar