← Search

Georg Rehm

8 accepted papers

2024

A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages

COLING 2024main

We present and describe two language resources in this paper: CATalog 1.0, the largest text corpus in Catalan to date, and CURATE (Corpus Utility for RAting TExt), a modular, parallelizable pipeline used for processing and scoring documents based on text quality that we have optimised to run in High…

2024

Common European Language Data Space

COLING 2024main

The Common European Language Data Space (LDS) is an integral part of the EU data strategy, which aims at developing a single market for data. Its decentralised technical infrastructure and governance scheme are currently being developed by the LDS project, which also has dedicated tasks for proof-of…

Cited by 2SourcePDFScholar
2024

European Language Grid: One Year after

COLING 2024main

The European Language Grid (ELG) is a cloud platform for the whole European Language Technology community. While the EU project that developed the platform successfully concluded in June 2022, the ELG initiative has continued. This article provides a description of the current state of ELG in terms…

Cited by 1SourcePDFScholar
2024

FoRC4CL: A Fine-grained Field of Research Classification and Annotated Dataset of NLP Articles

COLING 2024main

The steep increase in the number of scholarly publications has given rise to various digital repositories, libraries and knowledge graphs aimed to capture, manage, and preserve scientific data. Efficiently navigating such databases requires a system able to classify scholarly documents according to…

2024

Symmetric Dot-Product Attention for Efficient Training of BERT Language Models

ACL 2024findings

Initially introduced as a machine translation model, the Transformer architecture has now become the foundation for modern deep learning architecture, with applications in a wide range of fields, from computer vision to natural language processing. Nowadays, to tackle increasingly more complex tasks…

Cited by 2SourcePDFScholar
2022

HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information

ACL 2022findings

Transformer-based language models usually treat texts as linear sequences. However, most texts also have an inherent hierarchical structure, i.e., parts of a text can be identified using their position in this hierarchy. In addition, section titles usually indicate the common topic of their respecti…

2022

Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings

EMNLP 2022main

Learning scientific document representations can be substantially improved through contrastive learning objectives, where the challenge lies in creating positive and negative training samples that encode the desired similarity semantics. Prior work relies on discrete citation relations to generate c…

2020

Aspect-based Document Similarity for Research Papers

COLING 2020main

Traditional document similarity measures provide a coarse-grained distinction between similar and dissimilar documents. Typically, they do not consider in what aspects two documents are similar. This limits the granularity of applications like recommender systems that rely on document similarity. In…