← Search

Haoqi Yang

2 accepted papers

2025

XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks. However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in r

2024

Utilizing Second-Order Information in Noisy Information-Sharing Environments for Distributed Optimization

ICASSP 2024accepted

Decentralized optimization aims to cooperatively solve a global finite-sum loss function, where each agent only possesses knowledge of its own local function. Real-world applications introduce challenges such as unstable channels and differential privacy concerns, necessitating the development of mo…

Cited by 0SourceScholar