← Search

Fangcheng Shi

1 accepted papers

2026

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

ICML 2026poster

Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such as math and code. However, identifying an optimal mixture remains an open challenge, as existing approaches either rely o…

Cited by 0SourceScholar