2026
TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
AAAI 2026technical
Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little