← Search

Amir Zeldes

11 accepted papers

2024

DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing

COLING 2024main

This paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about…

2024

GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains

EMNLP 2024main

Work on shallow discourse parsing in English has focused on the Wall Street Journal corpus, the only large-scale dataset for the language in the PDTB framework. However, the data is not openly available, is restricted to the news domain, and is by now 35 years old. In this paper, we present and eval…

2024

SPLICE: A Singleton-Enhanced PipeLIne for Coreference REsolution

COLING 2024main

Singleton mentions, i.e. entities mentioned only once in a text, are important to how humans understand discourse from a theoretical perspective. However previous attempts to incorporate their detection in end-to-end neural coreference resolution for English have been hampered by the lack of singlet…

2024

To Ask LLMs about English Grammaticality, Prompt Them in a Different Language

EMNLP 2024finding

In addition to asking questions about facts in the world, some internet users—in particular, second language learners—ask questions about language itself. Depending on their proficiency level and audience, they may pose these questions in an L1 (first language) or an L2 (second language). We investi…

Cited by 2SourcePDFScholar
2024

UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies

COLING 2024main

The Universal Dependencies (UD) project has created an invaluable collection of treebanks with contributions in over 140 languages. However, the UD annotations do not tell the full story. Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elemen…

2024

Universal Anaphora: The First Three Years

COLING 2024main

The aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotati…

2023

ELQA: A Corpus of Metalinguistic Questions and Answers about English

ACL 2023long

We present ELQA, a corpus of questions and answers in and about the English language. Collected from two online forums, the >70k questions (from English learners and others) cover wide-ranging topics including grammar, meaning, fluency, and etymology. The answers include descriptions of general prop…

2023

GUMSum: Multi-Genre Data and Evaluation for English Abstractive Summarization

ACL 2023findings

Automatic summarization with pre-trained language models has led to impressively fluent results, but is prone to ‘hallucinations’, low performance on non-news genres, and outputs which are not exactly summaries. Targeting ACL 2023’s ‘Reality Check’ theme, we present GUMSum, a small but carefully cra…

2022

A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing

EMNLP 2022main

Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima’an et al. 2001). However, the data in HTB, a single-source newswire corpus, is now over 30 years old, and does not cover many aspects of contemporary Heb…

2021

OntoGUM: Evaluating Contextualized SOTA Coreference Resolution on 12 More Genres

ACL 2021short

SOTA coreference resolution produces increasingly impressive scores on the OntoNotes benchmark. However lack of comparable data following the same scheme for more genres makes it difficult to evaluate generalizability to open domain data. This paper provides a dataset and comprehensive evaluation sh…