← Search

Teresa Zhang

1 accepted papers

2026

Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling

AAAI 2026technical

Scaling long-context and agentic LLMs is increasingly limited by memory capacity and bandwidth rather than FLOPs. I propose an algorithmic framework for context engineering that models placement, compression, and scheduling as coupled optimization problems with explicit accuracy-efficiency trade-off

Cited by 0SourcePDFScholar