← Search

Keith Achorn

2 accepted papers

2021

Multilingual Spoken Words Corpus

NeurIPS 2021poster

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 3…

Cited by 66SourceScholar
2021

The People’s Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

NeurIPS 2021poster

The People’s Speech is a free-to-download 31,400-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA. The data is collected via searching the Internet for appropriately licensed audio data with existing transcription…

Cited by 99SourceScholar