← Search

Michael Y. Hu

3 accepted papers

2025

Aioli: A Unified Optimization Framework for Language Model Data Mixing

ICLR 2025poster

Language model performance depends on identifying the optimal mixture of data groups to train on (e.g., law, code, math). Prior work has proposed a diverse set of methods to efficiently learn mixture proportions, ranging from fitting regression models over training runs to dynamically updating propo…

2025

Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases

ACL 2025long

Pretraining language models on formal language can improve their acquisition of natural language. Which features of the formal language impart an inductive bias that leads to effective transfer? Drawing on insights from linguistics and complexity theory, we hypothesize that effective transfer occurs…

Cited by 0SourcePDFScholar
2025

Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check

EMNLP 2025

Downstream scaling laws aim to predict task performance at larger scales from the model’s performance at smaller scales. Whether such prediction should be possible is unclear: some works discover clear linear scaling trends after simple transformations of the performance metric, whereas others point