← Search

Ankit Yadav

2 accepted papers

2024

PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs

EMNLP 2024finding

Driven by the surge in code generation using large language models (LLMs), numerous benchmarks have emerged to evaluate these LLMs capabilities. We conducted a large-scale human evaluation of *HumanEval* and *MBPP*, two popular benchmarks for Python code generation, analyzing their diversity and dif…

2024

Remember This Event That Year? Assessing Temporal Information and Understanding in Large Language Models

EMNLP 2024finding

Large Language Models (LLMs) are increasingly ubiquitous, yet their ability to retain and reason about temporal information remains limited, hindering their application in real-world scenarios where understanding the sequential nature of events is crucial. Our study experiments with 12 state-of-the-…