← Search

Pu Yang

4 accepted papers

2026

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling

ICML 2026poster

While Mixture-of-Experts (MoE) architectures substantially bolster the expressive power of large-language models, their prohibitive memory footprint severely impedes the practical deployment on resource-constrained edge devices, especially when model behavior must be preserved without relying on los…

Cited by 0SourceScholar
2025

Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification

ICLR 2025poster

Large Language Models (LLM) are increasingly trained on data generated by other LLMs, either because generated text and images become part of the pre-training corpus, or because synthetized data is used as a replacement for expensive human-annotation. This raises concerns about *model collapse*, a d…

Cited by 3SourcePDFScholar
2025

Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping

NeurIPS 2025spotlight

Modern foundation models often undergo iterative ``bootstrapping'' in their post-training phase: a model generates synthetic data, an external verifier filters out low-quality samples, and the high-quality subset is used for further fine-tuning. Over multiple iterations, the model performance improv…

Cited by 0SourceScholar
2024

A Tale of Tails: Model Collapse as a Change of Scaling Laws

ICML 2024poster

As AI model size grows, neural *scaling laws* have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data. Yet, the widespread use of popular models means that the ecosystem of online data and text will co-…

Cited by 58SourcePDFScholar