← Search

Hendra Setiawan

4 accepted papers

2025

Beyond Text Compression: Evaluating Tokenizers Across Scales

ACL 2025long

The choice of tokenizer can profoundly impact language model performance, yet accessible and reliable evaluations of tokenizer quality remain an open challenge. Inspired by scaling consistency, we show that smaller models can accurately predict significant differences in tokenizer impact on larger m…

Cited by 0SourcePDFScholar
2023

Joint Speech Transcription and Translation: Pseudo-Labeling with Out-of-Distribution Data

ACL 2023findings

Self-training has been shown to be helpful in addressing data scarcity for many domains, including vision, speech, and language. Specifically, self-training, or pseudo-labeling, labels unsupervised data and adds that to the training pool. In this work, we investigate and use pseudo-labeling for a re…

Cited by 6SourcePDFScholar
2022

End-to-End Speech Translation for Code Switched Speech

ACL 2022findings

Code switching (CS) refers to the phenomenon of interchangeably using words and phrases from different languages. CS can pose significant accuracy challenges to NLP, due to the often monolingual nature of the underlying systems. In this work, we focus on CS in the context of English/Spanish conversa…