← Search

Zili Shao

2 accepted papers

2025

SmartCache: Context-aware Semantic Cache for Efficient Multi-turn LLM Inference

NeurIPS 2025poster

Large Language Models (LLMs) for multi-turn conversations suffer from inefficiency: semantically similar queries across different user sessions trigger redundant computation and duplicate memory-intensive Key-Value (KV) caches. Existing optimizations such as prefix caching overlook semantic similari…

Cited by 0SourceScholar
2018

A Reliable Video Storage Architecture in Hybrid SLC/MLC Nand Flash

ICASSP 2018accepted

In this paper, we propose a reliable video storage architecture in hybrid SLC/MLC storage systems. In this architecture, the video stream is reconstructed as the key cluster and the non-key cluster according to the importance of video restoration. The key cluster is stored in SLC blocks to ensure th…

Cited by 0SourceScholar