← Search

Mao-xun Huang

1 accepted papers

2025

Efficient Beam Search for Large Language Models Using Trie-Based Decoding

EMNLP 2025

This work presents a novel trie (prefix-tree)-based parallel decoding method that addresses the memory inefficiency of batch-based beam search. By sharing a single KV cache across beams with common prefixes, our approach dramatically reduces memory usage and enables efficient decoding. We evaluated