Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic Systems
Integrating Large Language Models (LLMs) into modern robotic systems presents significant computational and energy constraint challenges, particularly for human-centered robotic applications. This paper presents a novel hardware optimization technique for deploying LLMs on resource-constrained embed