← Search

Katarzyna Lorenc

2 accepted papers

2025

Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse

ACL 2025long

The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. We introduce forePLay, a novel Polish-language dataset for erotic content detection, comprising over 24,000 annotated sen…

2025

PLLuM-Align: Polish Preference Dataset for Large Language Model Alignment

EMNLP 2025

Alignment is the critical process of minimizing harmful outputs by teaching large language models (LLMs) to prefer safe, helpful and appropriate responses. While the majority of alignment research and datasets remain overwhelmingly English-centric, ensuring safety across diverse linguistic and cultu

Cited by 0SourcePDFScholar