← Search

Hengzhong Rao

1 accepted papers

2026

AlphaRouter: Token-level Routing Between SLM and LLM with Reinforcement Learning and Tree Search

ICML 2026poster

SLM-LLM routing accelerates generation by strategically invoking LLMs for critical tokens. However, existing methods typically train routers to mimic the LLM, capping performance at the reference trajectory's limit. In this work, we demonstrate that the SLM-LLM collaborative inference space offers a…

Cited by 0SourceScholar