← Search

Jeffrey Jian Ma

1 accepted papers

2025

Understanding Silent Data Corruption in LLM Training

ACL 2025long

As the scale of training large language models (LLMs) increases, one emergent failure is silent data corruption (SDC), where hardware produces incorrect computations without explicit failure signals. In this work, we are the first to investigate the impact of real-world SDCs on LLM training by compa…