2026
MOAI: Module-Optimizing Architecture for Non-Interactive Secure Transformer Inference
ICLR 2026poster
Privacy concerns have been raised in Large Language Models (LLM) inference when models are deployed in Cloud Service Providers (CSP). Homomorphic encryption (HE) offers a promising solution by enabling secure inference directly over encrypted inputs. However, the high computational overhead of HE re…