2025
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
NeurIPS 2025poster
Meeting service-level objectives (SLOs) in Large Language Models (LLMs) serving is critical, but managing the high variability in load presents a significant challenge. Recent advancements in FP8 inference, backed by native hardware support, offer a potential solution: executing FP16 models by defau…