← Search

Luoxin Chen

3 accepted papers

2025

AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training

EMNLP 2025

We introduce AdamS, a simple yet effective alternative to Adam for large language model (LLM) pretraining and post-training. By leveraging a novel denominator, i.e., the root of weighted sum of squares of the momentum and the current gradient, AdamS eliminates the need for second-moment estimates. H

2021

Industry Scale Semi-Supervised Learning for Natural Language Understanding

NAACL 2021industry

This paper presents a production Semi-Supervised Learning (SSL) pipeline based on the student-teacher framework, which leverages millions of unlabeled examples to improve Natural Language Understanding (NLU) tasks. We investigate two questions related to the use of unlabeled data in production SSL c…

Cited by 65SourcePDFScholar