2026
FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing
ICML 2026poster
Two-server secure inference allows a client to query a hosted large language model (LLM) without revealing prompts or embeddings. Recent GPU systems based on function secret sharing (FSS) make linear layers efficient, but fixed-point nonlinearities and helper operations remain a bottleneck because e…