← Search

Daniel Whitenack

2 accepted papers

2022

Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

EMNLP 2022main

We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. These datasets represent either the most, or among the most, multilingual datasets for each of the included d…

Cited by 23SourcePDFScholar
2022

Phone-ing it in: Towards Flexible Multi-Modal Language Model Training by Phonetic Representations of Data

ACL 2022long

Multi-modal techniques offer significant untapped potential to unlock improved NLP technology for local languages. However, many advances in language model pre-training are focused on text, a fact that only increases systematic inequalities in the performance of NLP tasks across the world’s language…