Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability. To overcome this challenge, we present prima.cpp, a distributed on-device inference system that runs 30-70B LLMs on consumer home clus…