2026
Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling
AAAI 2026technical
Scaling long-context and agentic LLMs is increasingly limited by memory capacity and bandwidth rather than FLOPs. I propose an algorithmic framework for context engineering that models placement, compression, and scheduling as coupled optimization problems with explicit accuracy-efficiency trade-off