← Search

Pingzhi Tang

4 accepted papers

2026

TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration

ICML 2026poster

Diffusion large language models (dLLMs) have recently gained significant attention due to their inherent support for parallel decoding. Building on this paradigm, Mixture-of-Experts (MoE) dLLMs with autoregressive (AR) initialization have further demonstrated strong performance competitive with main…

Cited by 0SourceScholar
2025

HD-PiSSA: High-Rank Distributed Orthogonal Adaptation

EMNLP 2025

Existing parameter-efficient fine-tuning (PEFT) methods for large language models (LLMs), such as LoRA and PiSSA, constrain model updates to low-rank subspaces, limiting their expressiveness and leading to suboptimal performance on complex tasks. To address this, we introduce **H**igh-rank **D**istr

Cited by 0SourcePDFScholar
2025

TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and Speedup

NeurIPS 2025spotlight

Modern large-language models often face communication bottlenecks on current hardware rather than computational limitations. *Multi-head latent attention (MLA)* addresses this by compressing the key-value cache using low-rank matrices, while the Absorb operation prevents the KV cache from reverting…

Cited by 0SourceScholar