← Search

Minghui Yu

3 accepted papers

2025

HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing

ICLR 2025poster

The frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leadin…

2025

SALS: Sparse Attention in Latent Space for KV Cache Compression

NeurIPS 2025poster

Large Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics withi…

Cited by 0SourceScholar
2020

Self-Prediction for Joint Instance and Semantic Segmentation of Point Clouds

ECCV 2020poster

We develop a novel learning scheme named Self-Prediction for 3D instance and semantic segmentation of point clouds. Distinct from most existing methods that focus on designing convolutional operators, our method designs a new learning scheme to enhance point relation exploring for better segmentatio…

Cited by 31SourcePDFScholar