← Search

Helen Jin

4 accepted papers

2025

Adaptively profiling models with task elicitation

EMNLP 2025

Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a method that automatically builds new evaluations to profile model behavior. Task elicitation finds hundreds of natural-lang

2025

Probabilistic Soundness Guarantees in LLM Reasoning Chains

EMNLP 2025

In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion. Current LLM-based error detection methods often fail to detect propagated errors because earlier errors can corrupt judgments of downstream reasoning.

2023

Generic Temporal Reasoning with Differential Analysis and Explanation

ACL 2023long

Temporal reasoning is the task of predicting temporal relations of event pairs. While temporal reasoning models can perform reasonably well on in-domain benchmarks, we have little idea of these systems’ generalizability due to existing datasets’ limitations. In this work, we introduce a novel task n…

Cited by 18SourcePDFScholar