← Search

Qihui Zhou

2 accepted papers

2026

TileSparse: Arithmetic-Intensity-Aware Sparse Attention for Compute-Bound LLM Decoding

ICML 2026poster

Sparse attention has emerged as a vital technique for long-context inference in Large Language Models (LLMs), effectively accelerating memory-bound decoding by reducing memory access for non-essential keys. However, the assumption that decoding attention is memory-bound has been shattered. The proli…

Cited by 0SourceScholar
2025

MixHD: A Method for Detecting Hallucinations Based on the Internal State and Output Probability of Large Language Models

ICASSP 2025accepted

This paper presents a novel hallucination detection method based on the internal states and output probabilities of large language models (LLMs) to address the common issue of hallucinations in model-generated content. We designed a new detection framework that extracts internal features such as hid…

Cited by 0SourceScholar