← Search

Bingsong Bai

4 accepted papers

2026

DEEP DUBBING: END-TO-END AUTO-AUDIOBOOK SYSTEM WITH TEXT-TO-TIMBRE AND CONTEXT-AWARE INSTRUCT-TTS

ICASSP 2026poster

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP models, whereas character voice timbre selection still relie…

Cited by 0SourcePDFScholar
2026

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

ICASSP 2026poster

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from incomplete speech, inaccurate or missing timestamps, and limited r…

Cited by 0SourcePDFScholar
2025

DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech

ICASSP 2025accepted

Traditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often…

Cited by 0SourceScholar