← Search

Nearchos Potamitis

2 accepted papers

2025

Cache Saver: A Modular Framework for Efficient, Affordable, and Reproducible LLM Inference

EMNLP 2025

Inference constitutes the majority of costs throughout the lifecycle of a large language model (LLM). While numerous LLM inference engines focusing primarily on low-level optimizations have been developed, there is a scarcity of non-intrusive client-side frameworks that perform high-level optimizati

2025

Fleet of Agents: Coordinated Problem Solving with Large Language Models

ICML 2025poster

While numerous frameworks have been developed to enhance the reasoning abilities of large language models (LLMs), there is a scarcity of methods that effectively balance the trade-off between cost and quality. In this paper, we introduce Fleet of Agents (FoA), a novel and intuitive yet principled f…