← Search

Kevin Ren

4 accepted papers

2026

QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

ICML 2026poster

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate …

Cited by 0SourceScholar
2025

Predicting Language Models’ Success at Zero-Shot Probabilistic Prediction

EMNLP 2025

Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics (e.g., to serve as risk models or augment survey datasets). However, when should a user have confidence that an LLM will provide high-quality predictions

2025

Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large-Scale Imitation Learning for MAPF

ICRA 2025

Multi-Agent Path Finding (MAPF) is the problem of effectively finding efficient collision-free paths for a group of agents in a shared workspace. The MAPF community has largely focused on developing high-performance heuristic search methods. Recently, several works have applied various machine learn

Cited by 7SourceScholar