← Search

Marina Neseem

3 accepted papers

2026

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

ICLR 2026oral

The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key–value (KV) cache, quickly overwhelming GPU memory. To address this challenge, we propose ThinKV, a thought-adaptive KV cache compression framework. ThinKV is b…

Cited by 12SourceScholar
2024

MTLoRA: Low-Rank Adaptation Approach for Efficient Multi-Task Learning

CVPR 2024highlight

Adapting models pre-trained on large-scale datasets to a variety of downstream tasks is a common strategy in deep learning. Consequently parameter-efficient fine-tuning methods have emerged as a promising way to adapt pre-trained models to different tasks while training only a minimal number of para…

2024

PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks

CVPR 2024poster

Low-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions batch normalization and quantization scaling dominate the inference cost o…

Cited by 1SourcePDFScholar