← Search

Hanzhi Zhang

2 accepted papers

2025

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

ACL 2025finding

Long-context understanding is crucial for many NLP applications, yet transformers struggle with efficiency due to the quadratic complexity of self-attention. Sparse attention methods alleviate this cost but often impose static, predefined masks, failing to capture heterogeneous attention patterns. T…

2023

Self-Distillation Hashing for Efficient Hamming Space Retrieval

ICASSP 2023accepted

Deep hashing-based approaches have become the optimal solutions for large-scale image retrieval task due to their high computational efficiency and low storage burden. Some methods leverage a large teacher network to improve the retrieval performance of the small student network through knowledge di…

Cited by 0SourceScholar