2026
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
ICML 2026poster
Large language models (LLMs) demand substantial computational and memory resources, posing challenges for efficient deployment. Two complementary approaches have emerged to address these issues: token-adaptive layer execution, which reduces floating-point operations (FLOPs) by selectively bypassing …