← Search

Hyunjik Jo

1 accepted papers

2025

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information

IJCAI 2025

How can we accelerate large language models (LLMs) without sacrificing accuracy? The slow inference speed of LLMs hinders us to benefit from their remarkable performance in diverse applications. This is mainly because numerous sublayers are stacked together in LLMs. Sublayer pruning compresses and e