← Search

Sicheng Lai

2 accepted papers

2026

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

ICML 2026poster

Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmar…

Cited by 0SourceScholar
2025

Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM

EMNLP 2025

The rapid advancement of multimodal large language models (MLLMs) has significantly enhanced performance across benchmarks. However, data contamination — partial/entire benchmark data is included in the model’s training set — poses critical challenges for fair evaluation. Existing detection methods