← Search

Abhishek Arora

6 accepted papers

2024

Contrastive Entity Coreference and Disambiguation for Historical Texts

EMNLP 2024main

Massive-scale historical document collections are crucial for social science research. Despite increasing digitization, these documents typically lack unique cross-document identifiers for individuals mentioned within the texts, as well as individual identifiers from external knowledge bases like Wi…

2024

LinkTransformer: A Unified Package for Record Linkage with Transformer Language Models

ACL 2024system demonstrations

Many computational analyses require linking information across noisy text datasets. While large language models (LLMs) offer significant promise, approximate string matching packages in popular statistical softwares such as R and Stata remain predominant in academic applications. These packages have…

2024

Newswire: A Large-Scale Structured Database of a Century of Historical News

NeurIPS 2024poster

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a national identity and shared understanding of the world, but there is no comprehensive archive of the content sent over…

2023

American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers

NeurIPS 2023poster

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other layout regions. OCR quality can also be low. This study deve…

Cited by 29SourcePDFScholar
2023

Quantifying Character Similarity with Vision Transformers

EMNLP 2023long main

Record linkage is a bedrock of quantitative social science, as analyses often require linking data from multiple, noisy sources. Off-the-shelf string matching methods are widely used, as they are straightforward and cheap to implement and scale. Not all character substitutions are equally probable,…

Cited by 0SourcecodeScholar