← Search

Jiaqi Song

2 accepted papers

2026

RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression

ICML 2026poster

Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks that limit expressiveness under heterogeneous data geometry. While recent dynamic quantizers like QINCo adapt codebooks to individual inputs and improve…

Cited by 0SourceScholar
2024

FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model

EMNLP 2024industry

In this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction. Unlike other modalities such as vision or text, speech has unique temporal dependencies, making previous efficient inference works on other modalities not directly applicable. Further…