← Search

Darshan Deshpande

2 accepted papers

2026

Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis

ICML 2026poster

Recent advances in reinforcement learning for code generation have made robust environments essential to prevent reward hacking. As LLMs increasingly serve as evaluators in code-based RL, their ability to detect reward hacking remains understudied. In this paper, we propose a novel taxonomy of rewar…

Cited by 0SourceScholar
2024

Contextualizing Argument Quality Assessment with Relevant Knowledge

NAACL 2024short

Automatic assessment of the quality of arguments has been recognized as a challenging task with significant implications for misinformation and targeted speech. While real-world arguments are tightly anchored in context, existing computational methods analyze their quality in isolation, which affect…