← Search

Mu Yuan

3 accepted papers

2025

A-VL: Adaptive Attention for Large Vision-Language Models

AAAI 2025technical

The Large Vision-Language Model (LVLM) integrates computer vision and natural language processing techniques, offering substantial application potential. However, these models demand extensive resources during inference. Adaptive attention techniques can dynamically reduce computational redundancy a…

2025

RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service

ACL 2025finding

Retrieval-augmented generation (RAG) improves the service quality of large language models by retrieving relevant documents from credible literature and integrating them into the context of the user query.Recently, the rise of the cloud RAG service has made it possible for users to query relevant do…

2022

MLink: Linking Black-Box Models for Collaborative Multi-Model Inference

AAAI 2022technical

The cost efficiency of model inference is critical to real-world machine learning (ML) applications, especially for delay-sensitive tasks and resource-limited devices. A typical dilemma is: in order to provide complex intelligent services (e.g. smart city), we need inference results of multiple ML m…