2026
Residual Tokens Enhance Masked Autoencoders for Speech Modeling
ICASSP 2026oral
Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments the supervised attributes-based modeling with unsupervised re…