2026
Stratos: An End-to-End Distillation Pipeline for Customized LLMs Under Distributed Cloud Environments
AAAI 2026technical
The growing industrial demand for customized and cost-efficient large language models (LLMs) is fueled by the rise of vertical, domain-specific tasks and the need to optimize performance under constraints such as latency and budget. Knowledge distillation, as an efficient model compression and trans