← Search

Qijun Miao

2 accepted papers

2026

Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMs

ICML 2026poster

The rapid advancement of large language models (LLMs) has led to remarkable performance across diverse domains, making them indispensable assistants in daily life and work. Currently, LLM services are primarily accessed in two ways: (i) paid access to cloud-hosted LLMs, which are powerful but introd…

Cited by 0SourceScholar
2025

User-side Model Consistency Monitoring for Open Source Large Language Models Inference Services

ACL 2025long

With the continuous advancement in the performance of open-source large language models (LLMs), their inference services have attracted a substantial user base by offering quality comparable to closed-source models at a significantly lower cost. However, it has also given rise to trust issues regard…