← Search

Chris Parnin

3 accepted papers

2026

jqBench: a benchmark for reading and editing JSON from natural language and/or examples

ICLR 2026poster

We introduce jqBench, a new benchmark for evaluating language models on JSON querying and transformation tasks, where the intent can be given specified using natural language and/or examples. Whereas jqBench is mainly aimed at using the `jq` tool, it can be used to evaluate other programming languag…

Cited by 0SourceScholar
2025

Evaluating the Evaluator: Measuring LLMs’ Adherence to Task Evaluation Instructions

AAAI 2025technical

LLMs-as-a-judge is a recently popularized method which replaces human judgements in task evaluation with automatic evaluation using LLMs. Due to widespread use of RLHF (Reinforcement Learning from Human Feedback), state-of-the-art LLMs like GPT4 and Llama3 are expected to have strong alignment with…

Cited by 10SourcePDFScholar
2025

TeCoFeS: Text Column Featurization using Semantic Analysis

NAACL 2025findings

Extracting insights from text columns can bechallenging and time-intensive. Existing methods for topic modeling and feature extractionare based on syntactic features and often overlook the semantics. We introduce the semantictext column featurization problem, and presenta scalable approach for autom…

Cited by 0SourcePDFScholar