← Search

Deuk Sin Kwon

2 accepted papers

2022

BECEL: Benchmark for Consistency Evaluation of Language Models

COLING 2022main

Behavioural consistency is a critical condition for a language model (LM) to become trustworthy like humans. Despite its importance, however, there is little consensus on the definition of LM consistency, resulting in different definitions across many studies. In this paper, we first propose the ide…

2022

KoBEST: Korean Balanced Evaluation of Significant Tasks

COLING 2022main

A well-formulated benchmark plays a critical role in spurring advancements in the natural language processing (NLP) field, as it allows objective and precise evaluation of diverse models. As modern language models (LMs) have become more elaborate and sophisticated, more difficult benchmarks that req…