← Search

Heshaam Faili

5 accepted papers

2025

IRUEX: A Study on Large Language Models Problem-Solving Skills in Iran’s University Entrance Exam

COLING 2025main

In this paper, we present the IRUEX dataset, a novel multiple-choice educational resource specifically designed to evaluate the performance of Large Language Models (LLMs) across seven distinct categories. The dataset contains 868 Iran university entrance exam questions (Konkour) and 36,485 addition…

Cited by 0SourcePDFScholar
2025

Matina: A Large-Scale 73B Token Persian Text Corpus

NAACL 2025long

Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian has often been underrepresented due to limited resources fo…

2024

EPOQUE: An English-Persian Quality Estimation Dataset

COLING 2024main

Translation quality estimation (QE) is an important component in real-world machine translation applications. Unfortunately, human labeled QE datasets, which play an important role in developing and assessing QE models, are only available for limited language pairs. In this paper, we present the fir…

2024

Esposito: An English-Persian Scientific Parallel Corpus for Machine Translation

COLING 2024main

Neural machine translation requires large number of parallel sentences along with in-domain parallel data to attain best results. Nevertheless, no scientific parallel corpus for English-Persian language pair is available. In this paper, a parallel corpus called Esposito is introduced, which contains…

Cited by 1SourcePDFScholar
2023

PMI-Align: Word Alignment With Point-Wise Mutual Information Without Requiring Parallel Training Data

ACL 2023findings

Word alignment has many applications including cross-lingual annotation projection, bilingual lexicon extraction, and the evaluation or analysis of translation outputs. Recent studies show that using contextualized embeddings from pre-trained multilingual language models could give us high quality w…