← Search

Kaheer Suleman

6 accepted papers

2024

Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective

ACL 2024findings

It is increasingly common to evaluate the same coreference resolution (CR) model on multiple datasets. Do these multi-dataset evaluations allow us to draw meaningful conclusions about model generalization? Or, do they rather reflect the idiosyncrasies of a particular experimental setup (e.g., the sp…

2023

The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources

ACL 2023long

Many state-of-the-art natural language understanding (NLU) models are based on pretrained neural language models. These models often make inferences using information from multiple sources. An important class of such inferences are those that require both background knowledge, presumably contained i…

2022

Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

NAACL 2022long

There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess varying quality criteria depending on the deployment setting. While the landscape of NLG evaluation has been well-mappe…

Cited by 36SourcePDFScholar
2021

ADEPT: An Adjective-Dependent Plausibility Task

ACL 2021long

A false contract is more likely to be rejected than a contract is, yet a false key is less likely than a key to open doors. While correctly interpreting and assessing the effects of such adjective-noun pairs (e.g., false key) on the plausibility of given events (e.g., opening doors) underpins many n…

2021

Modeling Event Plausibility with Consistent Conceptual Abstraction

NAACL 2021long

Understanding natural language requires common sense, one aspect of which is the ability to discern the plausibility of events. While distributional models—most recently pre-trained, Transformer language models—have demonstrated improvements in modeling event plausibility, their performance still fa…

2020

An Analysis of Dataset Overlap on Winograd-Style Tasks

COLING 2020main

The Winograd Schema Challenge (WSC) and variants inspired by it have become important benchmarks for common-sense reasoning (CSR). Model performance on the WSC has quickly progressed from chance-level to near-human using neural language models trained on massive corpora. In this paper, we analyze th…