2026
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
ICASSP 2026poster
Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored phenomenon undermines the coherence of synthetic speech, especially…