← Search

Feiyang Ren

2 accepted papers

2026

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

ICML 2026poster

Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top-$p$ nucleus sampling) offer superior accuracy by dynamically fluctuating memory budgets, yet modern inference engines (e.g., vLLM) demand rigid, static memory patterns to leverage CUDA Graphs and…

Cited by 0SourceScholar
2025

Abstain Mask Retain Core: Time Series Prediction by Adaptive Masking Loss with Representation Consistency

NeurIPS 2025spotlight

Time series forecasting plays a pivotal role in critical domains such as energy management and financial markets. Although deep learning-based approaches (e.g., MLP, RNN, Transformer) have achieved remarkable progress, the prevailing "long-sequence information gain hypothesis" exhibits inherent limi…

Cited by 0SourcecodeScholar