← Search

Junqing Lin

2 accepted papers

2026

CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints

AAAI 2026technical

Mixture of Experts (MoE) models have emerged as a promising approach to scale language models efficiently by activating only a subset of parameters for each input. However, deploying these models under GPU memory constraints remains challenging, as existing offloading strategies incur significant ov

Cited by 0SourcePDFScholar
2025

Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their extensive parameter scales pose significant challenges for practical deployment. Unstructured pruning has emerged as an effective model compression strategy with minimal performance loss, which introduces fine-grained…

Cited by 0SourceScholar