← Search

Shuaihong Jiang

3 accepted papers

2026

GateRA: Token-aware Modulation for Parameter-Efficient Fine-tuning

AAAI 2026technical

Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, DoRA, and HiRA, enable lightweight adaptation of large pre-trained models via low-rank updates. However, existing PEFT approaches apply static, input-agnostic updates to all tokens, disregarding the varying importance and difficulty of d

Cited by 0SourcePDFScholar
2025

Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps

ACL 2025finding

Retrieval-augmented generation (RAG) has emerged as a pivotal method for expanding the knowledge of large language models. To handle complex queries more effectively, researchers developed Adaptive-RAG (A-RAG) to enhance the generated quality through multiple interactions with external knowledge bas…

2025

Enhancing Large Language Model Inference Efficiency via Lookahead Cache Filtering

ICASSP 2025accepted

The large Key-Value (KV) cache is a significant challenge in deploying Large Language Models (LLMs). Current research addressing these issues employs cache compression techniques, which we find suffer from information loss and the "lost-in-the-middle" problem. We propose the Lookahead Cache Filterin…

Cited by 0SourceScholar