← Search

Nick McKenna

4 accepted papers

2025

Evaluating the Evaluator: Measuring LLMs’ Adherence to Task Evaluation Instructions

AAAI 2025technical

LLMs-as-a-judge is a recently popularized method which replaces human judgements in task evaluation with automatic evaluation using LLMs. Due to widespread use of RLHF (Reinforcement Learning from Human Feedback), state-of-the-art LLMs like GPT4 and Llama3 are expected to have strong alignment with…

Cited by 10SourcePDFScholar
2025

Unlocking SLM Potential for Data Analysis Code Generation via Non-Parametric Knowledge Distillation

NeurIPS 2025poster

Knowledge distillation from Large Language Models (LLMs) to locally hosted Small Language Models (SLMs) provides advantages for Data Analysis Code Generation (DACG) such as privacy protection. However, achieving effective distillation without resource-intensive training is challenging. This paper in…

Cited by 0SourceScholar
2023

Sources of Hallucination by Large Language Models on Inference Tasks

EMNLP 2023long findings

Large Language Models (LLMs) are claimed to be capable of Natural Language Inference (NLI), necessary for applied tasks like question answering and summarization. We present a series of behavioral studies on several LLM families (LLaMA, GPT-3.5, and PaLM) which probe their behavior using controlled…

Cited by 0SourcecodeScholar
2021

Multivalent Entailment Graphs for Question Answering

EMNLP 2021main

Drawing inferences between open-domain natural language predicates is a necessity for true language understanding. There has been much progress in unsupervised learning of entailment graphs for this purpose. We make three contributions: (1) we reinterpret the Distributional Inclusion Hypothesis to m…

Cited by 17SourcePDFScholar