← Search

Tingwei Guo

5 accepted papers

2025

Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models

EMNLP 2025

End-to-end Large Speech Language Models (LSLMs) have demonstrated impressive conversational generation abilities, yet consistently fall short of traditional pipeline systems on semantic understanding benchmarks. In this work, we reveal through systematic experimentation that although LSLMs lose some

Cited by 0SourcePDFScholar
2022

Audio Deepfake Detection System with Neural Stitching for ADD 2022

ICASSP 2022accepted

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge[1]. The very same system was used for both two rounds of evaluation in Track 3.2 with similar training methodology. The first round of Track 3.2 data is generated from Text-to-Speech…

Cited by 0SourceScholar
2022

Audio-Visual Wake Word Spotting System for MISP Challenge 2021

ICASSP 2022accepted

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the environmental robustness of far-field wake word spotting. In the propose…

Cited by 0SourceScholar
2022

Time Domain Adversarial Voice Conversion for ADD 2022

ICASSP 2022accepted

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language content into target speaker’s fake speech. Then the converted spee…

Cited by 0SourceScholar
2021

Didispeech: A Large Scale Mandarin Speech Corpus

ICASSP 2021accepted

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is recorded in quiet environment and is suitable for various speech p…

Cited by 0SourceScholar