← Search

Zhiping Xiu

5 accepted papers

2025

Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation

ICASSP 2025accepted

Large language models (LLMs) have revolutionized natural language processing (NLP) with impressive performance across various text-based tasks. However, the extension of text-dominant LLMs to with speech generation tasks remains underexplored. In this work, we introduce a text-to-speech (TTS) system…

Cited by 6SourceScholar
2025

Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

ICASSP 2025accepted

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech translation model that outputs discrete speech tokens in a low…

Cited by 0SourceScholar
2024

Ultra-Lightweight Neural Differential DSP Vocoder for High Quality Speech Synthesis

ICASSP 2024accepted

Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing (DSP) based vocoder can be implemented via lightweight fast Fou…

Cited by 5SourceScholar
2022

Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation Complexity

ICASSP 2022accepted

Low bitrate speech codecs have become an area of intense research. Traditional speech codecs, which use signal processing methods to encode and decode speech, often suffer from quality issues at low bitrates. A neural speech codec, which uses a deep neural network in the compression pipeline, can he…

Cited by 0SourceScholar
2021

Multi-Rate Attention Architecture for Fast Streamable Text-to-Speech Spectrum Modeling

ICASSP 2021accepted

Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates the actual audio. High-quality spectrum models usually incorporate the encoder-decoder architecture with self-attention…

Cited by 0SourceScholar