← Search

Michael Frank

2 accepted papers

2024

DevBench: A multimodal developmental benchmark for language learning

NeurIPS 2024oral

How (dis)similar are the learning trajectories of vision–language models and children? Recent modeling work has attempted to understand the gap between models’ and humans’ data efficiency by constructing models trained on less data, especially multimodal naturalistic data. However, such models are o…

2024

Is Child-Directed Speech Effective Training Data for Language Models?

EMNLP 2024main

While high-performing language models are typically trained on hundreds of billions of words, human children become fluent language users with a much smaller amount of data. What are the features of the data they receive, and how do these features support language modeling objectives? To investigate…