2025
From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement
EMNLP 2025
Chain-of-Thought (CoT) reasoning improves performance on complex tasks but introduces significant inference latency due to its verbosity. In this work, we propose Multiround Adaptive Chain-of-Thought Compression ( MACC ), a framework that leverages the token elasticity phenomenon —where overly small