← Search

Tyler McDonald

3 accepted papers

2025

Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index

COLING 2025main

As prompt engineering research rapidly evolves, evaluations beyond accuracy are crucial for developing cost-effective techniques. We present the Economical Prompting Index (EPI), a novel metric that combines accuracy scores with token consumption, adjusted by a user-specified cost concern level to r…

2025

NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers

COLING 2025main

Large Language Models (LLMs) have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable. We present NYT-Connections, a collection of 358 simple word classification puzzles derived from the New York Times Connections game. This be…

Cited by 1SourcePDFScholar
2024

STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions

EMNLP 2024main

Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies evaluate scenarios in isolation, without considering the broader context or the spectrum of potential biases within eac…