← Search

David Kanter

5 accepted papers

2023

DataPerf: Benchmarks for Data-Centric AI Development

NeurIPS 2023poster

Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and…

2022

The Dollar Street Dataset: Images Representing the Geographic and Socioeconomic Diversity of the World

NeurIPS 2022accept

It is crucial that image datasets for computer vision are representative and contain accurate demographic information to ensure their robustness and fairness, especially for smaller subpopulations. To address this issue, we present Dollar Street - a supervised dataset that contains 38,479 images of…

Cited by 79SourcePDFScholar
2021

MLPerf Tiny Benchmark

NeurIPS 2021poster

Advancements in ultra-low-power tiny machine learning (TinyML) systems promise to unlock an entirely new class of smart applications. However, continued progress is limited by the lack of a widely accepted and easily reproducible benchmark for these systems. To meet this need, we present MLPerf Tiny…

Cited by 246SourceScholar
2021

Multilingual Spoken Words Corpus

NeurIPS 2021poster

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 3…

Cited by 66SourceScholar
2021

The People’s Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

NeurIPS 2021poster

The People’s Speech is a free-to-download 31,400-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA. The data is collected via searching the Internet for appropriately licensed audio data with existing transcription…

Cited by 99SourceScholar