← Search

Zhiyuan Tang

10 accepted papers

2026

CHAIN OF CORRECTION FOR FULL-TEXT SPEECH RECOGNITION WITH LARGE LANGUAGE MODELS

ICASSP 2026poster

Full-text error correction with Large Language Models (LLMs) for Automatic Speech Recognition (ASR) is attracting increased attention for its ability to address a wide range of error types, such as punctuation restoration and inverse text normalization, across long context. However, challenges remai…

Cited by 0SourcePDFScholar
2025

Full-text Error Correction for Chinese Speech Recognition with Large Language Model

ICASSP 2025accepted

Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper i…

Cited by 0SourceScholar
2021

KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects

NeurIPS 2021poster

This paper introduces an open source speech dataset, KeSpeech, which involves 1,542 hours of speech signals recorded by 27,237 speakers in 34 cities in China, and the pronunciation includes standard Mandarin and its 8 subdialects. The new dataset possesses several properties. Firstly, the dataset pr…

Cited by 41SourceScholar
2018

Human and Machine Speaker Recognition Based on Short Trivial Events

ICASSP 2018accepted

Human speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. How…

Cited by 0SourceScholar
2017

Memory visualization for gated recurrent neural networks in speech recognition

ICASSP 2017accepted

Recurrent neural networks (RNNs) have shown clear superiority in sequence modeling, particularly the ones with gated units, such as long short-term memory (LSTM) and gated recurrent unit (GRU). However, the dynamic properties behind the remarkable performance remain unclear in many applications, e.g…

Cited by 0SourceScholar