← Search

David Piorkowski

2 accepted papers

2025

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

NeurIPS 2025poster

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different domains. However, finding suitable benchmarks is difficult gi…

Cited by 0SourcecodeScholar
2024

Language Models in Dialogue: Conversational Maxims for Human-AI Interactions

EMNLP 2024finding

Modern language models, while sophisticated, exhibit some inherent shortcomings, particularly in conversational settings. We claim that many of the observed shortcomings can be attributed to violation of one or more conversational principles. By drawing upon extensive research from both the social s…

Cited by 12SourcePDFScholar