FRONTEND TOKEN ENHANCEMENT FOR TOKEN-BASED SPEECH RECOGNITION
Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these representations, such as semantic or phonetic tokens derived from clustering output…