← Search

Jiyang Zhang

2 accepted papers

2026

PLSemanticsBench: A Formal Semantics Reasoning Benchmark for Code

ICML 2026poster

Recent work asks whether large language models (LLMs) condition their reasoning on explicit rules rather than statistical regularities from pretraining. Program execution provides a canonical instance: formal semantics define behavior through sym- bolic transition rules that can be systematically al…

Cited by 0SourceScholar
2022

Impact of Evaluation Methodologies on Code Summarization

ACL 2022long

There has been a growing interest in developing machine learning (ML) models for code summarization tasks, e.g., comment generation and method naming. Despite substantial increase in the effectiveness of ML models, the evaluation methodologies, i.e., the way people split datasets into training, vali…