← Search

Krzysztof Jurkiewicz

2 accepted papers

2025

Oddballness: universal anomaly detection with language models

COLING 2025main

We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but instead of focusing on low-likelihood tokens, it considers a new m…

2022

Challenging America: Modeling language in longer time scales

NAACL 2022findings

The aim of the paper is to apply, for historical texts, the methodology used commonly to solve various NLP tasks defined for contemporary data, i.e. pre-train and fine-tune large Transformer models. This paper introduces an ML challenge, named Challenging America (ChallAm), based on OCR-ed excerpts…

Cited by 3SourcePDFScholar