← Search

Akshat Ramachandran

2 accepted papers

2026

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

ICLR 2026oral

The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key–value (KV) cache, quickly overwhelming GPU memory. To address this challenge, we propose ThinKV, a thought-adaptive KV cache compression framework. ThinKV is b…

Cited by 0SourceScholar
2025

OuroMamba: A Data-Free Quantization Framework for Vision Mamba

ICCV 2025poster

We present OuroMamba, the first data-free post-training quantization (DFQ) method for vision Mamba-based models (VMMs). We identify two key challenges in enabling DFQ for VMMs, (1) VMM's recurrent state transitions restricts the capturing of long-range interactions and leads to semantically weak syn…