2024
StackEval: Benchmarking LLMs in Coding Assistance
NeurIPS 2024poster
We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated datasets: StackEval, a large-scale benchmark derived from Stack O…