Automatic Partitioning of a Code-Switched Speech Corpus Using Mixed-Integer Programming
Defining training, development and test set partitions for speech corpora is usually accomplished by hand. However, for the dataset under investigation, which contains a large number of speakers, eight different languages and code-switching between all the languages, this style of partitioning is no…