2025
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
NeurIPS 2025poster
We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly for large language model (LLM) training with multiple GPUs and multiple SSDs. Its…