← Search

Ziran Qin

5 accepted papers

2026

Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling

AAAI 2026technical

Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of key-value (KV)

Cited by 0SourcePDFScholar
2026

Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers

ICLR 2026poster

Massive Activations (MAs) are a well-documented phenomenon across Transformer architectures, and prior studies in both LLMs and ViTs have shown that they play a substantial role in shaping model behavior. However, the nature and function of MAs within Diffusion Transformers (DiTs) remain largely une…

Cited by 0SourceScholar
2026

VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding

ICML 2026poster

Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this AR paradigm inevitably faces a dual efficiency bottleneck: strictly unidirectional attention compromises *understanding …

Cited by 3SourceScholar
2025

CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences

ICLR 2025poster

Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference burden, they often fail to allocate resources rationally across layers with different attention patterns. In this paper, w…

2025

Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver

ICASSP 2025accepted

Mathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems, which require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to understand geometric diagrams, performing no better in geo…

Cited by 0SourceScholar