← Search

Sheng Qi

3 accepted papers

2026

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

ICML 2026poster

Deploying multiple models within shared GPU clusters is a key strategy to improve resource efficiency in large language model (LLM) serving. Existing multi-LLM serving systems improve GPU utilization at the cost of degraded inference performance, particularly time-to-first-token (TTFT). We attribute…

Cited by 0SourceScholar
2025

ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents

ICLR 2025poster

Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, t…

2021

Multi-Target Invisibly Trojaned Networks for Visual Recognition and Detection

IJCAI 2021poster

Visual backdoor attack is a recently-emerging task which aims to implant trojans in a deep neural model. A trojaned model responds to a trojan-invoking trigger in a fully predictable manner while functioning normally otherwise. As a key motivating fact to this work, most triggers adopted in existing…

Cited by 4SourcePDFScholar