← Search

Sarik Ghazarian

6 accepted papers

2024

AXCEL: Automated eXplainable Consistency Evaluation using LLMs

EMNLP 2024finding

Large Language Models (LLMs) are widely used in both industry and academia for various tasks, yet evaluating the consistency of generated text responses continues to be a challenge. Traditional metrics like ROUGE and BLEU show a weak correlation with human judgment. More sophisticated metrics using…

2023

ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue Systems

ACL 2023long

Commonsense reasoning is omnipresent in human communications and thus is an important feature for open-domain dialogue systems. However, evaluating commonsense in dialogue systems is still an open challenge. We take the first step by focusing on event commonsense that considers events and their rela…

2022

DEAM: Dialogue Coherence Evaluation using AMR-based Semantic Manipulations

ACL 2022long

Automatic evaluation metrics are essential for the rapid development of open-domain dialogue systems as they facilitate hyper-parameter tuning and comparison between models. Although recently proposed trainable conversation-level metrics have shown encouraging results, the quality of the metrics is…

2022

What is wrong with you?: Leveraging User Sentiment for Automatic Dialog Evaluation

ACL 2022findings

Accurate automatic evaluation metrics for open-domain dialogs are in high demand. Existing model-based metrics for system response evaluation are trained on human annotated data, which is cumbersome to collect. In this work, we propose to use information that can be automatically extracted from the…

2021

DiSCoL: Toward Engaging Dialogue Systems through Conversational Line Guided Response Generation

NAACL 2021system demonstrations

Having engaging and informative conversations with users is the utmost goal for open-domain conversational systems. Recent advances in transformer-based language models and their applications to dialogue systems have succeeded to generate fluent and human-like responses. However, they still lack con…

Cited by 14SourcePDFScholar
2021

Plot-guided Adversarial Example Construction for Evaluating Open-domain Story Generation

NAACL 2021long

With the recent advances of open-domain story generation, the lack of reliable automatic evaluation metrics becomes an increasingly imperative issue that hinders the fast development of story generation. According to conducted researches in this regard, learnable evaluation metrics have promised mor…