← Search

Ali Emami

18 accepted papers

2025

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models

EMNLP 2025

Research on bias in Text-to-Image (T2I) models has primarily focused on demographic representation and stereotypical attributes, overlooking a fundamental question: how does grammatical gender influence visual representation across languages? We introduce a cross-linguistic benchmark examining words

Cited by 0SourcePDFScholar
2025

Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index

COLING 2025main

As prompt engineering research rapidly evolves, evaluations beyond accuracy are crucial for developing cost-effective techniques. We present the Economical Prompting Index (EPI), a novel metric that combines accuracy scores with token consumption, adjusted by a user-specified cost concern level to r…

2025

Fine-Tuned LLMs are “Time Capsules” for Tracking Societal Bias Through Books

NAACL 2025long

Books, while often rich in cultural insights, can also mirror societal biases of their eras—biases that Large Language Models (LLMs) may learn and perpetuate during training. We introduce a novel method to trace and quantify these biases using fine-tuned LLMs. We develop BookPAGE, a corpus comprisin…

2025

NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers

COLING 2025main

Large Language Models (LLMs) have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable. We present NYT-Connections, a collection of 358 simple word classification puzzles derived from the New York Times Connections game. This be…

Cited by 1SourcePDFScholar
2025

Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks

EMNLP 2025

As Large Language Models (LLMs) increasingly integrate into everyday workflows, where users shape outcomes through multi-turn collaboration, a critical question emerges: do users with different personality traits systematically prefer certain LLMs over others? We conduc-ted a study with 32 participa

Cited by 0SourcePDFScholar
2025

Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations

ACL 2025finding

Addressing gender bias and maintaining logical coherence in machine translation remains challenging, particularly when translating between natural gender languages, like English, and genderless languages, such as Persian, Indonesian, and Finnish. We introduce the Translate-with-Care (TWC) dataset, c…

2025

We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

EMNLP 2025

Large language models (LLMs) struggle to navigate culturally specific communication norms, limiting their effectiveness in global contexts. We focus on Persian *taarof*, a social norm in Iranian interactions, which is a sophisticated system of ritual politeness that emphasizes deference, modesty, an

Cited by 0SourcePDFScholar
2024

Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models

ACL 2024long

As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models. We introduce the concept of Confidence-Probability Alignment…

2024

MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models

EMNLP 2024main

This study explores the effectiveness of Large Language Models (LLMs) in creating personalized “mirror stories” that reflect and resonate with individual readers’ identities, addressing the significant lack of diversity in literature. We present MirrorStories, a corpus of 1,500 personalized short st…

Cited by 1SourcePDFScholar
2024

Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge

ACL 2024long

Large Language Models (LLMs) have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning. However, applying this reasoning to multimodal domains, where understanding text and images together is essential, remains a substa…

2024

STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions

EMNLP 2024main

Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies evaluate scenarios in isolation, without considering the broader context or the spectrum of potential biases within eac…

2024

Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models

ACL 2024long

Research on Large Language Models (LLMs) has often neglected subtle biases that, although less apparent, can significantly influence the models’ outputs toward particular social narratives. This study addresses two such biases within LLMs: representative bias, which denotes a tendency of LLMs to gen…

2023

Debiasing should be Good and Bad: Measuring the Consistency of Debiasing Techniques in Language Models

ACL 2023findings

Debiasing methods that seek to mitigate the tendency of Language Models (LMs) to occasionally output toxic or inappropriate text have recently gained traction. In this paper, we propose a standardized protocol which distinguishes methods that yield not only desirable results, but are also consistent…

2021

ADEPT: An Adjective-Dependent Plausibility Task

ACL 2021long

A false contract is more likely to be rejected than a contract is, yet a false key is less likely than a key to open doors. While correctly interpreting and assessing the effects of such adjective-noun pairs (e.g., false key) on the plausibility of given events (e.g., opening doors) underpins many n…

2020

An Analysis of Dataset Overlap on Winograd-Style Tasks

COLING 2020main

The Winograd Schema Challenge (WSC) and variants inspired by it have become important benchmarks for common-sense reasoning (CSR). Model performance on the WSC has quickly progressed from chance-level to near-human using neural language models trained on massive corpora. In this paper, we analyze th…

2019

Exploiting Uncertainty of Deep Neural Networks for Improving Segmentation Accuracy in MRI Images

ICASSP 2019accepted

Deep neural networks have shown great achievements in solving complex problems. However, there are fundamental challenges which limit their real world applications. Lack of a measurable criterion for estimating uncertainty of the network predictions is one of these challenges. However, we can comput…

Cited by 0SourceScholar