2025
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
ACL 2025long
As large language models (LLMs) scale in size and adoption, their computational and environmental costs continue to rise. Prior benchmarking efforts have primarily focused on latency reduction in idealized settings, often overlooking the diverse real-world inference workloads that shape energy use.…