← Search

Omin Kwon

1 accepted papers

2025

NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs

NeurIPS 2025poster

Meeting service-level objectives (SLOs) in Large Language Models (LLMs) serving is critical, but managing the high variability in load presents a significant challenge. Recent advancements in FP8 inference, backed by native hardware support, offer a potential solution: executing FP16 models by defau…

Cited by 0SourcecodeScholar