2024
AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data
ICASSP 2024accepted
Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that…