← Search

Farhan Samir

5 accepted papers

2025

ZIPA: A family of efficient models for multilingual phone recognition

ACL 2025long

We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPA PACK++, a large-scale multilingual speech corpus with 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing u…

2024

Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia

EMNLP 2024main

To explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses. These studies often rely on coarse corpus-level statistics or local word-level analyses, mainly in English. We introduce the InfoGap method—an efficient and…

2024

The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language

NAACL 2024long

In this project, we demonstrate that phoneme-based models for speech processing can achieve strong crosslinguistic generalizability to unseen languages. We curated the IPAPACK, a massively multilingual speech corpora with phonemic transcriptions, encompassing more than 115 languages from diverse lan…

2023

Understanding Compositional Data Augmentation in Typologically Diverse Morphological Inflection

EMNLP 2023long main

Data augmentation techniques are widely used in low-resource automatic morphological inflection to address the issue of data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we aim to shed light on the theoretical aspects of the data augmentation…

Cited by 0SourcecodeScholar
2022

Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Under-Documented Languages

ACL 2022findings

Recent progress in NLP is driven by pretrained models leveraging massive datasets and has predominantly benefited the world’s political and economic superpowers. Technologically underserved languages are left behind because they lack such resources. Hundreds of underserved languages, nevertheless, h…