OPTIMIZING SPEECH LANGUAGE MODELS FOR ACOUSTIC CONSISTENCY
We study speech language models that incorporate semantic initialization and planning losses to achieve robust and consistent generation. Our approach initializes speech tokens with self-supervised features, applies a light alignment loss, and trains with thinning and auxiliary objectives that targe…