← Search

Hsinyu Tsai

2 accepted papers

2026

Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees

ICLR 2026poster

Sparse Mixture-of-Experts (MoE) allows scaling of language and vision models efficiently by activating only a small subset of experts per input. While this reduces computation, the large number of parameters still incurs substantial memory overhead during inference. Post-training quantization has be…

Cited by 0SourcecodeScholar
2025

Analog Foundation Models

NeurIPS 2025poster

Analog in-memory computing (AIMC) is a promising compute paradigm to improve speed and power efficiency of neural network inference beyond the limits of conventional von Neumann-based architectures. However, AIMC introduces fundamental challenges such as noisy computations and strict constraints on…

Cited by 0SourcecodeScholar