← Search

Colin Leong

4 accepted papers

2023

JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing

EMNLP 2023long findings

Advancements in sign language processing have been hindered by a lack of sufficient data, impeding progress in recognition, translation, and production tasks. The absence of comprehensive sign language datasets across the world's sign languages has widened the gap in this field, resulting in a few s…

Cited by 0SourcecodeScholar
2022

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

NAACL 2022long

Recent advances in the pre-training for language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these datasets. This is primarily because many widely spoken languages that are not well represented on the web and therefore ex…

2022

Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

EMNLP 2022main

We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. These datasets represent either the most, or among the most, multilingual datasets for each of the included d…

Cited by 23SourcePDFScholar
2022

Phone-ing it in: Towards Flexible Multi-Modal Language Model Training by Phonetic Representations of Data

ACL 2022long

Multi-modal techniques offer significant untapped potential to unlock improved NLP technology for local languages. However, many advances in language model pre-training are focused on text, a fact that only increases systematic inequalities in the performance of NLP tasks across the world’s language…