← Search

Roberto Bifulco

2 accepted papers

2025

What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering

NAACL 2025long

Large Language Models (LLMs) changed the way we design and interact with software systems. Their ability to process and extract information from text has drastically improved productivity in a number of routine tasks. Developers that want to include these models in their software stack, however, fac…

2024

AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents

NAACL 2024system demonstrations

The advances made by Large Language Models (LLMs) have led to the pursuit of LLM agents that can solve intricate, multi-step reasoning tasks. As with any research pursuit, benchmarking and evaluation are key corner stones to efficient and reliable progress. However, existing benchmarks are often nar…