← Search

Francesco Ortu

4 accepted papers

2025

Language Model Alignment in Multilingual Trolley Problems

ICLR 2025spotlight

We evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200+ countries, we develop a cross-lingual corpus of moral dilemma vignettes in ove…

Cited by 3SourcePDFScholar
2025

Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing

ACL 2025finding

The ability of Natural Language Processing (NLP) methods to categorize text into multiple classes has motivated their use in online content moderation tasks, such as hate speech and fake news detection. However, there is limited understanding of how or why these methods make such decisions, or why c…

2025

The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models

NeurIPS 2025poster

Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-language models (VLMs) handle image-understanding tasks, focusing on how visual information is processed and transferred…

Cited by 0SourceScholar
2024

Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals

ACL 2024long

Interpretability research aims to bridge the gap between the empirical success and our scientific understanding of the inner workings of large language models (LLMs). However, most existing research in this area focused on analyzing a single mechanism, such as how models copy or recall factual knowl…