← Search

Oskar Holmström

2 accepted papers

2024

How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?

EMNLP 2024finding

Work on instruction-tuned Large Language Models (LLMs) has used automatic methods based on text overlap and LLM judgments as cost-effective alternatives to human evaluation. In this paper, we perform a meta-evaluation of such methods and assess their reliability across a broad range of tasks. In eva…

2024

Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark

ACL 2024short

Text-to-SQL, which involves translating natural language into Structured Query Language (SQL), is crucial for enabling broad access to structured databases without expert knowledge. However, designing models for such tasks is challenging due to numerous factors, including the presence of noise, such…