LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
As large language models (LLMs) scale, their inference incurs substantial computational resources, exposing them to energy-latency attacks, where crafted prompts induce high energy and latency cost. Existing attack methods aim to prolong output by delaying the generation of termination symbols. Howe