← Search

Naman Garg

3 accepted papers

2025

REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites

NeurIPS 2025poster

We introduce REAL, a benchmark and framework for multi-turn agent evaluations on deterministic simulations of real-world websites. REAL comprises high-fidelity, deterministic replicas of 11 widely-used websites across domains such as e-commerce, travel, communication, and professional networking. We…

Cited by 0SourceScholar
2024

RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

NeurIPS 2024poster

Training on model-generated synthetic data is a promising approach for finetuning LLMs, but it remains unclear when it helps or hurts. In this paper, we investigate this question for math reasoning via an empirical study, followed by building a conceptual understanding of our observations. First, we…

2024

Recursive Introspection: Teaching Language Model Agents How to Self-Improve

NeurIPS 2024poster

A central piece in enabling intelligent agentic behavior in foundation models is to make them capable of introspecting upon their behavior, reasoning, and correcting their mistakes as more computation or interaction is available. Even the strongest proprietary large language models (LLMs) do not qui…

Cited by 39SourcePDFScholar