← Search

Ehsaneddin Asgari

7 accepted papers

2025

Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation

ACL 2025finding

Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data. Retrieval-Augmented Generation (RAG) mitigates these issues by integrating external dynamic information for improved factual grounding. With advances in multimodal learning,…

2025

Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description

NAACL 2025findings

3D facial emotion modeling has important applications in areas such as animation design, virtual reality, and emotional human-computer interaction (HCI). However, existing models are constrained by limited emotion classes and insufficient datasets. To address this, we introduce Emo3D, an extensive “…

Cited by 0SourcePDFScholar
2025

Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages

NAACL 2025short

While broad-coverage multilingual natural language processing tools have been developed, a significant portion of the world’s over 7000 languages are still neglected. One reason is the lack of evaluation datasets that cover a diverse range of languages, particularly those that are low-resource or en…

2024

The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments

COLING 2024main

While human values play a crucial role in making arguments persuasive, we currently lack the necessary extensive datasets to develop methods for analyzing the values underlying these arguments on a large scale. To address this gap, we present the Touché23-ValueEval dataset, an expansion of the Webis…

2024

Transformers for Bridging Persian Dialects: Transliteration Model for Tajiki and Iranian Scripts

COLING 2024main

In this study, we address the linguistic challenges posed by Tajiki Persian, a distinct variant of the Persian language that utilizes the Cyrillic script due to historical “Russification”. This distinguishes it from other Persian dialects that adopt the Arabic script. Despite its profound linguistic…

2024

TuringQ: Benchmarking AI Comprehension in Theory of Computation

EMNLP 2024finding

We present TuringQ, the first benchmark designed to evaluate the reasoning capabilities of large language models (LLMs) in the theory of computation. TuringQ consists of 4,006 undergraduate and graduate-level question-answer pairs, categorized into four difficulty levels and covering seven core theo…

2021

KnowMAN: Weakly Supervised Multinomial Adversarial Networks

EMNLP 2021main

The absence of labeled data for training neural models is often addressed by leveraging knowledge about the specific task, resulting in heuristic but noisy labels. The knowledge is captured in labeling functions, which detect certain regularities or patterns in the training samples and annotate corr…