← Search

Michael Gertz

5 accepted papers

2024

LexDrafter: Terminology Drafting for Legislative Documents Using Retrieval Augmented Generation

COLING 2024main

With the increase in legislative documents at the EU, the number of new terms and their definitions is increasing as well. As per the Joint Practical Guide of the European Parliament, the Council and the Commission, terms used in legal documents shall be consistent, and identical concepts shall be e…

2024

Numbers Matter! Bringing Quantity-awareness to Retrieval Systems

EMNLP 2024finding

Quantitative information plays a crucial role in understanding and interpreting the content of documents. Many user queries contain quantities and cannot be resolved without understanding their semantics, e.g., “car that costs less than $10k”. Yet, modern search engines apply the same ranking mechan…

2022

EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain

EMNLP 2022main

Existing summarization datasets come with two main drawbacks: (1) They tend to focus on overly exposed domains, such as news articles or wiki-like texts, and (2) are primarily monolingual, with few multilingual datasets.In this work, we propose a novel dataset, called EUR-Lex-Sum, based on manually…

2022

Three Real-World Datasets and Neural Computational Models for Classification Tasks in Patent Landscaping

EMNLP 2022main

Patent Landscaping, one of the central tasks of intellectual property management, includes selecting and grouping patents according to user-defined technical or application-oriented criteria. While recent transformer-based models have been shown to be effective for classifying patents into taxonomie…