← Search

Ruohua Zhou

2 accepted papers

2026

DEEP DUBBING: END-TO-END AUTO-AUDIOBOOK SYSTEM WITH TEXT-TO-TIMBRE AND CONTEXT-AWARE INSTRUCT-TTS

ICASSP 2026poster

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP models, whereas character voice timbre selection still relie…

Cited by 0SourcePDFScholar
2026

POLY-SVC: POLYPHONY-AWARE SINGING VOICE CONVERSION WITH HARMONIC MODELING

ICASSP 2026poster

Singing Voice Conversion (SVC) aims to transform a source singing voice into a target singer while preserving lyrics and melody. Most existing SVC methods depend on F0 extractors to capture the lead melody from clean vocals. However, no existing method can reliably extract clean vocals from accompan…

Cited by 0SourcePDFScholar