← Search

Xiaomeng Han

2 accepted papers

2026

NLI : Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference

ICLR 2026poster

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks, but their deployment is often constrained by substantial memory footprints and computational costs. While prior work has achieved significant progress in compressing and accelerating linear layers, no…

Cited by 0SourceScholar
2025

Pushing the Limits of BFP on Narrow Precision LLM Inference

AAAI 2025technical

The substantial computational and memory demands of Large Language Models (LLMs) hinder their deployment. Block Floating Point (BFP) has proven effective in accelerating linear operations, a cornerstone of LLM workloads. However, as sequence lengths grow, nonlinear operations, such as Attention, in…

Cited by 0SourcePDFScholar