DRPBench: Evaluating LLMs in Concurrent Code Comprehension via Fine-grained Data Race Prediction
Large Language Models (LLMs) have demonstrated sophisticated comprehension of sequential code, yet their capacity for reasoning about concurrent programs remains largely unquantified. We introduce DRPBench, a benchmark designed to evaluate the concurrent code comprehension of LLMs by measuring their…