← Search

Tarek Naous

9 accepted papers

2026

Flipping the Dialogue: Training and Evaluating User Language Models

ICLR 2026poster

Conversations with LMs involve two participants: a human user leading the conversation, and an LM assistant responding to the user's request. To satisfy this specific role, LMs are post-trained to be helpful assistants - optimized to produce exhaustive and well-structured responses, often free of am…

Cited by 0SourceScholar
2025

CARE: Multilingual Human Preference Learning for Cultural Awareness

EMNLP 2025

Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied. In this paper, we systematically analyze how native human cultural preferences can be incorpora

2025

On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

NAACL 2025long

Language Models (LMs) have been shown to exhibit a strong preference towards entities associated with Western culture when operating in non-Western languages. In this paper, we aim to uncover the origins of entity-related cultural biases in LMs by analyzing several contributing factors, including th…

2024

Having Beer after Prayer? Measuring Cultural Bias in Large Language Models

ACL 2024long

As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial. Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances. In this paper, we show that multilingual and Arabic monolin…

2024

ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment

EMNLP 2024main

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This paper introduces ReadMe++, a multilingual multi-domain data…

2024

Reducing Privacy Risks in Online Self-Disclosures with Language Models

ACL 2024long

Self-disclosure, while being common and rewarding in social media interaction, also poses privacy risks. In this paper, we take the initiative to protect the user-side privacy associated with online self-disclosure through detection and abstraction. We develop a taxonomy of 19 self-disclosure catego…

Cited by 19SourcePDFScholar
2023

Revisiting non-English Text Simplification: A Unified Multilingual Benchmark

ACL 2023long

Recent advancements in high-quality, large-scale English resources have pushed the frontier of English Automatic Text Simplification (ATS) research. However, less work has been done on multilingual text simplification due to the lack of a diverse evaluation benchmark that covers complex-simple sente…

2022

Stanceosaurus: Classifying Stance Towards Multicultural Misinformation

EMNLP 2022main

We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi and Arabic annotated with stance towards 250 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims. The claims in Stanceosaurus originate from 15 fact-check…

Cited by 18SourcePDFScholar