← Search

Emilio Ferrara

9 accepted papers

2026

Position: Measuring Human Preferences in RLHF is a Social Science Problem

ICML 2026spotlight

RLHF assumes that annotation responses reflect genuine human preferences. We argue this assumption warrants systematic examination, and that behavioral science offers frameworks that bring clarity to when it holds and when it breaks down. Behavioral scientists have documented for sixty years that pe…

Cited by 0SourceScholar
2025

Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment

EMNLP 2025

As LLMs become central to interactive applications, ranging from tutoring to mental health, the ability to express personality in culturally appropriate ways is increasingly important. While recent works have explored personality evaluation of LLMs, they largely overlook the interplay between cultur

2025

IOHunter: Graph Foundation Model to Uncover Online Information Operations

AAAI 2025technical

Social media platforms have become vital spaces for public discourse, serving as modern agorás where a wide range of voices influence societal narratives. However, their open nature also makes them vulnerable to exploitation by malicious actors, including state-sponsored entities, who can conduct in…

2024

Can Language Model Moderators Improve the Health of Online Discourse?

NAACL 2024long

Conversational moderation of online communities is crucial to maintaining civility for a constructive environment, but it is challenging to scale and harmful to moderators. The inclusion of sophisticated natural language generation modules as a force multiplier to aid human moderators is a tantalizi…

2024

Explaining Mixtures of Sources in News Articles

EMNLP 2024finding

Human writers plan, _then_ write. For large language models (LLMs) to play a role in longer-form article generation, we must understand the planning steps humans make before writing. We explore one kind of planning, source-selection in news, as a case-study for evaluating plans in long-form generati…

Cited by 2SourcePDFScholar
2024

Tracking the Newsworthiness of Public Documents

ACL 2024long

Journalists regularly make decisions on whether or not to report stories, based on “news values”. In this work, we wish to explicitly model these decisions to explore _when_ and _why_ certain stories get press attention. This is challenging because very few labelled links between source documents an…

2023

Controlled Text Generation with Hidden Representation Transformations

ACL 2023findings

We propose CHRT (Control HiddenRepresentation Transformation) – a con-trolled language generation framework thatsteers large language models to generatetext pertaining to certain attributes (such astoxicity). CHRT gains attribute control bymodifying the hidden representation of thebase model through…

2023

FairFed: Enabling Group Fairness in Federated Learning

AAAI 2023technical

Training ML models which are fair across different demographic groups is of critical importance due to the increased integration of ML in crucial decision-making scenarios such as healthcare and recruitment. Federated learning has been viewed as a promising solution for collaboratively training mach…

Cited by 237SourcePDFScholar
2023

Identifying Informational Sources in News Articles

EMNLP 2023long main

News articles are driven by the informational sources journalists use in reporting. Modeling when, how and why sources get used together in stories can help us better understand the information we consume and even help journalists with the task of producing it. In this work, we take steps toward thi…

Cited by 0SourcecodeScholar