← Search

Ziming Dai

2 accepted papers

2026

Stratos: An End-to-End Distillation Pipeline for Customized LLMs Under Distributed Cloud Environments

AAAI 2026technical

The growing industrial demand for customized and cost-efficient large language models (LLMs) is fueled by the rise of vertical, domain-specific tasks and the need to optimize performance under constraints such as latency and budget. Knowledge distillation, as an efficient model compression and trans

Cited by 0SourcePDFScholar
2026

TS-PEFT: Unveiling Token-Level Redundancy in Parameter-Efficient Fine-Tuning

IJCAI 2026

Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate under an implicit assumption: once a target module is selected, every token passing through it contributes equally to the downstream task and requires a parameter update. In this paper, we challenge this convention by revealing

Cited by 0Scholar