← Search

Albert Sawczyn

5 accepted papers

2025

Hallucination Detection in LLMs Using Spectral Features of Attention Maps

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable performance across various tasks but remain prone to hallucinations. Detecting hallucinations is essential for safety-critical applications, and recent methods leverage attention map properties to this end, though their effectiveness remains

2025

The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs

EMNLP 2025

Large language models (LLMs) have revolutionized natural language processing, yet their tendency to hallucinate poses serious challenges for reliable deployment. Despite numerous hallucination detection methods, their evaluations often rely on ROUGE, a metric based on lexical overlap that misaligns

Cited by 0SourcePDFScholar
2024

Developing PUGG for Polish: A Modern Approach to KBQA, MRC, and IR Dataset Construction

ACL 2024findings

Advancements in AI and natural language processing have revolutionized machine-human language interactions, with question answering (QA) systems playing a pivotal role. The knowledge base question answering (KBQA) task, utilizing structured knowledge graphs (KG), allows for handling extensive knowle…

2024

Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings

COLING 2024main

Knowledge-intensive tasks pose a significant challenge for Machine Learning (ML) techniques. Commonly adopted methods, such as Large Language Models (LLMs), often exhibit limitations when applied to such tasks. Nevertheless, there have been notable endeavours to mitigate these challenges, with a sig…

2022

This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish

NeurIPS 2022accept

The availability of compute and data to train larger and larger language models increases the demand for robust methods of benchmarking the true progress of LM training. Recent years witnessed significant progress in standardized benchmarking for English. Benchmarks such as GLUE, SuperGLUE, or KILT…