Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features…