2025
Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving
ICML 2025poster
In this paper, we investigate how to push the performance limits of serving Deep Neural Network (DNN) models on CPU-based servers. Specifically, we observe that while intra-operator parallelism across multiple threads is an effective way to reduce inference latency, it provides diminishing returns.…