← Search

Jonas Zausinger

1 accepted papers

2025

Regress, Don't Guess: A Regression-like Loss on Number Tokens for Language Models

ICML 2025poster

While language models have exceptional capabilities at text generation, they lack a natural inductive bias for emitting numbers and thus struggle in tasks involving quantitative reasoning, especially arithmetic. One fundamental limitation is the nature of the Cross Entropy loss, which assumes a nomi…