← Search

Yongliang Tao

2 accepted papers

2026

MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

CVPR 2026

Post-training quantization (PTQ) with computational equivalence for Large Language Models (LLMs) have demonstrated remarkable advances, however, their application to Multimodal Large Language Models (MLLMs) presents substantial challenges. In this paper, we analyze SmoothQuant as a case study and id

Cited by 0SourcecodeScholar
2025

La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation

ICML 2025poster

Activation sparsity can reduce the computational overhead and memory transfers during the forward pass of Large Language Model (LLM) inference. Existing methods face limitations, either demanding time-consuming recovery training that hinders real-world adoption, or relying on empirical magnitude-bas…

Cited by 0SourcePDFScholar