← Search

Garrett Tanzer

10 accepted papers

2025

FSboard: Over 3 Million Characters of ASL Fingerspelling Collected via Smartphones

CVPR 2025poster

Progress in machine understanding of sign languages has been slow and hampered by limited data. In this paper, we present FSboard, an American Sign Language fingerspelling dataset situated in a mobile text entry use case, collected from 147 paid and consenting Deaf signers using Pixel 4A selfie came…

Cited by 1SourcePDFScholar
2025

YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel Corpus

ICLR 2025poster

Even for better-studied sign languages like American Sign Language (ASL), data is the bottleneck for machine learning research. The situation is worse yet for the many other sign languages used by Deaf/Hard of Hearing communities around the world. In this paper, we present YouTube-SL-25, a large-sca…

2024

A Benchmark for Learning to Translate a New Language from One Grammar Book

ICLR 2024spotlight

Large language models (LLMs) can perform impressive feats with in-context learning or lightweight finetuning. It is natural to wonder how well these models adapt to genuinely new tasks, but how does one find tasks that are unseen in internet-scale training sets? We turn to a field that is explicitly…

Cited by 40SourcePDFScholar
2024

DOCCI: Descriptions of Connected and Contrasting Images

ECCV 2024poster

"Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would allow for richer associations to be learned by models. To fill the gap, we introduce Descriptions of Connected and Contra…

Cited by 52SourcePDFScholar
2024

Reconsidering Sentence-Level Sign Language Translation

EMNLP 2024main

Historically, sign language machine translation has been posed as a sentence-level task: datasets consisting of continuous narratives are chopped up and presented to the model as isolated clips. In this work, we explore the limitations of this task framing. First, we survey a number of linguistic ph…

Cited by 1SourcePDFScholar
2023

PopSign ASL v1.0: An Isolated American Sign Language Dataset Collected via Smartphones

NeurIPS 2023poster

PopSign is a smartphone-based bubble-shooter game that helps hearing parents of deaf infants learn sign language. To help parents practice their ability to sign, PopSign is integrating sign language recognition as part of its gameplay. For training the recognizer, we introduce the PopSign ASL v1.0 d…

Cited by 14SourcePDFScholar
2023

YouTube-ASL: A Large-Scale, Open-Domain American Sign Language-English Parallel Corpus

NeurIPS 2023poster

Machine learning for sign languages is bottlenecked by data. In this paper, we present YouTube-ASL, a large-scale, open-domain corpus of American Sign Language (ASL) videos and accompanying English captions drawn from YouTube. With ~1000 hours of videos and >2500 unique signers, YouTube-ASL is ~3x a…

Cited by 52SourcePDFScholar