← Search

Amin Ahmad

3 accepted papers

2025

FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs

NAACL 2025short

Summarization is one of the most common tasks performed by large language models (LLMs), especially in applications like Retrieval-Augmented Generation (RAG). However, existing evaluations of hallucinations in LLM-generated summaries, and evaluations of hallucination detection models both suffer fro…

2025

MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems

NAACL 2025long

Traditional retrieval-augmented generation (RAG) benchmarks evaluate systems using heuristic-based metrics, but these require human preferences as the ground truth for reference. In contrast, arena-based benchmarks, where systems compete against each other, require an expensive large language model…

2023

mAggretriever: A Simple yet Effective Approach to Zero-Shot Multilingual Dense Retrieval

EMNLP 2023short main

Multilingual information retrieval (MLIR) is a crucial yet challenging task due to the need for human annotations in multiple languages, making training data creation labor-intensive. In this paper, we introduce mAggretriever, which effectively leverages semantic and lexical features from pre-train…

Cited by 0SourceScholar