2026
Speech Recognition Model Improves Text-to-Speech Synthesis Using Fine-Grained Reward
AAAI 2026technical
Recent advancements in Text-to-Speech (TTS) technology have been remarkable, enabling current models to clone arbitrary unseen speakers and synthesize high-quality, natural-sounding speech. However, corresponding evaluation techniques appear to be lagging: existing Mean Opinion Score (MOS) estimatio