← Search

Zhiqing Zhong

2 accepted papers

2026

SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

ICML 2026poster

The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80\%. However, we reveal that this performance is inflated: our re-evaluation demonstrates that one in five "solved" patches from the top-30 agents are semantically incorrect, passing only because weak tes…

Cited by 0SourceScholar
2025

OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?

ICLR 2025poster

Large language models (LLMs) are driving substantial advancements in software engineering, with successful applications like Copilot and Cursor transforming real-world development practices. However, current research predominantly focuses on the early stages of development, such as code generation,…

Cited by 2SourcePDFScholar