← Search

Ruicheng Ao

4 accepted papers

2026

PILOT-Bench: Probabilistic Interaction for LLM Operations in Tool-driven Scenarios

ICLR 2026poster

We introduce PILOT-Bench, a benchmark that evaluates LLM workflow execution under simulated realistic conditions of instruction quality variability and tool execution uncertainty. Unlike existing benchmarks that encounter these challenges incidentally, our work makes uncertainty the primary focus of…

Cited by 0SourcecodeScholar
2026

Solver-in-the-Loop: MDP-Based Benchmarks for Self-Correction and Behavioral Rationality in Operations Research

ICML 2026poster

Operations Research practitioners routinely debug infeasible models through an iterative process: analyzing Irreducible Infeasible Subsystems (\IIS{}), identifying constraint conflicts, and systematically repairing formulations until feasibility is achieved. Yet existing LLM benchmarks evaluate OR a…

Cited by 0SourceScholar
2025

Learning to price with resource constraints: from full information to machine-learned prices

NeurIPS 2025poster

Dynamic pricing with resource constraints is a critical challenge in online learning, requiring a delicate balance between exploring unknown demand patterns and exploiting known information to maximize revenue. We propose three tailored algorithms to address this problem across varying levels of pri…

Cited by 0SourceScholar