← Search

Smaranda Muresan

32 accepted papers

2026

Death of the Novel(ty): Beyond N-Gram Novelty as a Metric for Textual Creativity

ICLR 2026poster

$N$-gram novelty is widely used to evaluate language models' ability to generate text outside of their training data. More recently, it has also been adopted as a metric for measuring textual creativity. However, theoretical work on creativity suggests that this approach may be inadequate, as it doe…

Cited by 0SourcecodeScholar
2026

LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

ICML 2026poster

Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating such systems remains challenging. We introduce LiveNewsBench, a rigorous and regularly updated benchmark designed to ass…

Cited by 0SourceScholar
2025

Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning

ACL 2025long

We introduce Browsing Lost Unformed Recollections, a tip-of-the-tongue known-item search and reasoning benchmark for general AI assistants. BLUR introduces a set of 573 real-world validated questions that demand searching and reasoning across multimodal and multilingual inputs, as well as proficient…

2025

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

EMNLP 2025

Large Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced human perspectives. Recent research has underscored the importance of enabling LLMs to support steerable pluralism — the c

Cited by 0SourcePDFScholar
2025

Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution

COLING 2025main

Recent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the la…

Cited by 0SourcePDFScholar
2025

Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layers

EMNLP 2025

We propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models. We evaluate our approach on two popular authorship attribution models and three evaluation datasets, in in-domain a

Cited by 0SourcePDFScholar
2025

Understanding Figurative Meaning through Explainable Visual Entailment

NAACL 2025long

Large Vision-Language Models (VLMs) have demonstrated strong capabilities in tasks requiring a fine-grained understanding of literal meaning in images and text, such as visual question-answering or visual entailment. However, there has been little exploration of the capabilities of these models when…

2024

Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game

EMNLP 2024main

The New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts. We collect438 Connections games to evaluate the performance of state-of-the-art large language models (LLMs) against expert and novice humanplayers. Our results show that even the best-p…

2024

ICLEF: In-Context Learning with Expert Feedback for Explainable Style Transfer

ACL 2024long

While state-of-the-art large language models (LLMs) can excel at adapting text from one style to another, current work does not address the explainability of style transfer models. Recent work has explored generating textual explanations from larger teacher models and distilling them into smaller st…

2024

Identifying Self-Disclosures of Use, Misuse and Addiction in Community-based Social Media Posts

NAACL 2024findings

In the last decade, the United States has lost more than 500,000 people from an overdose involving prescription and illicit opioids making it a national public health emergency (USDHHS, 2017). Medical practitioners require robust and timely tools that can effectively identify at-risk patients. Commu…

2024

Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition

ACL 2024findings

Recognizing fallacies is crucial for ensuring the quality and validity of arguments across various domains. However, computational fallacy recognition faces challenges due to the diverse genres, domains, and types of fallacies found in datasets. This leads to a highly multi-class, and even multi-lab…

2023

I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphors

ACL 2023findings

Visual metaphors are powerful rhetorical devices used to persuade or communicate creative ideas through images. Similar to linguistic metaphors, they convey meaning implicitly through symbolism and juxtaposition of the symbols. We propose a new task of generating visual metaphors from linguistic met…

2023

Learning to Follow Object-Centric Image Editing Instructions Faithfully

EMNLP 2023long findings

Natural language instructions are a powerful interface for editing the outputs of text-to-image diffusion models. However, several challenges need to be addressed: 1) underspecification (the need to model the implicit meaning of instructions) 2) grounding (the need to localize where the edit has to…

Cited by 0SourcecodeScholar
2023

NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-Fly

EMNLP 2023long main

Knowledge of norms is needed to understand and reason about acceptable behavior in human communication and interactions across sociocultural scenarios. Most computational research on norms has focused on a single culture, and manually built datasets, from non-conversational settings. We address thes…

Cited by 0SourcecodeScholar
2023

NormDial: A Comparable Bilingual Synthetic Dialog Dataset for Modeling Social Norm Adherence and Violation

EMNLP 2023short main

Social norms fundamentally shape interpersonal communication. We present NormDial, a high-quality dyadic dialogue dataset with turn-by-turn annotations of social norm adherences and violations for Chinese and American cultures. Introducing the task of social norm observance detection, our dataset is…

Cited by 0SourcecodeScholar
2023

Sociocultural Norm Similarities and Differences via Situational Alignment and Explainable Textual Entailment

EMNLP 2023long main

Designing systems that can reason across cultures requires that they are grounded in the norms of the contexts in which they operate. However, current research on developing computational models of social norms has primarily focused on American society. Here, we propose a novel approach to discover…

Cited by 0SourcecodeScholar
2022

CONSISTENT: Open-Ended Question Generation From News Articles

EMNLP 2022finding

Recent work on question generation has largely focused on factoid questions such as who, what,where, when about basic facts. Generating open-ended why, how, what, etc. questions thatrequire long-form answers have proven more difficult. To facilitate the generation of openended questions, we propose…

2022

FLUTE: Figurative Language Understanding through Textual Explanations

EMNLP 2022main

Figurative language understanding has been recently framed as a recognizing textual entailment (RTE) task (a.k.a. natural language inference (NLI)). However, similar to classical RTE/NLI datasets they suffer from spurious correlations and annotation artifacts. To tackle this problem, work on NLI has…

2022

Multitask Instruction-based Prompting for Fallacy Recognition

EMNLP 2022main

Fallacies are used as seemingly valid arguments to support a position and persuade the audience about its validity. Recognizing fallacies is an intrinsically difficult task both for humans and machines. Moreover, a big challenge for computational models lies in the fact that fallacies are formulated…

2022

Unsupervised Stem-based Cross-lingual Part-of-Speech Tagging for Morphologically Rich Low-Resource Languages

NAACL 2022long

Unsupervised cross-lingual projection for part-of-speech (POS) tagging relies on the use of parallel data to project POS tags from a source language for which a POS tagger is available onto a target language across word-level alignments. The projected tags then form the basis for learning a POS mode…

2021

COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 Pandemic

ACL 2021long

We introduce a FEVER-like dataset COVID-Fact of 4,086 claims concerning the COVID-19 pandemic. The dataset contains claims, evidence for the claims, and contradictory claims refuted by the evidence. Unlike previous approaches, we automatically detect true claims and their source articles and then ge…

2021

Don’t Go Far Off: An Empirical Study on Neural Poetry Translation

EMNLP 2021main

Despite constant improvements in machine translation quality, automatic poetry translation remains a challenging problem due to the lack of open-sourced parallel poetic corpora, and to the intrinsic complexities involved in preserving the semantics, style and figurative nature of poetry. We present…

2021

ENTRUST: Argument Reframing with Language Models and Entailment

NAACL 2021long

Framing involves the positive or negative presentation of an argument or issue depending on the audience and goal of the speaker. Differences in lexical framing, the focus of our work, can have large effects on peoples’ opinions and beliefs. To make progress towards reframing arguments for positive…

2021

Emotion-Infused Models for Explainable Psychological Stress Detection

NAACL 2021long

The problem of detecting psychological stress in online posts, and more broadly, of detecting people in distress or in need of help, is a sensitive application for which the ability to interpret models is vital. Here, we present work exploring the use of a semantically related task, emotion detectio…

2021

Implicit Premise Generation with Discourse-aware Commonsense Knowledge Models

EMNLP 2021main

Enthymemes are defined as arguments where a premise or conclusion is left implicit. We tackle the task of generating the implicit premise in an enthymeme, which requires not only an understanding of the stated conclusion and premise but also additional inferences that could depend on commonsense kno…

2021

MERMAID: Metaphor Generation with Symbolism and Discriminative Decoding

NAACL 2021long

Generating metaphors is a challenging task as it requires a proper understanding of abstract concepts, making connections between unrelated concepts, and deviating from the literal meaning. In this paper, we aim to generate a metaphoric sentence given a literal expression by replacing relevant verbs…

2021

Metaphor Generation with Conceptual Mappings

ACL 2021long

Generating metaphors is a difficult task as it requires understanding nuanced relationships between abstract concepts. In this paper, we aim to generate a metaphoric sentence given a literal expression by replacing relevant verbs. Guided by conceptual metaphor theory, we propose to control the gener…

2021

Weakly-Supervised Methods for Suicide Risk Assessment: Role of Related Domains

ACL 2021short

Social media has become a valuable resource for the study of suicidal ideation and the assessment of suicide risk. Among social media platforms, Reddit has emerged as the most promising one due to its anonymity and its focus on topic-based communities (subreddits) that can be indicative of someone’s…

2020

Fact vs. Opinion: the Role of Argumentation Features in News Classification

COLING 2020main

A 2018 study led by the Media Insight Project showed that most journalists think that a clearmarking of what is news reporting and what is commentary or opinion (e.g., editorial, op-ed)is essential for gaining public trust. We present an approach to classify news articles into newsstories (i.e., rep…

Cited by 27SourcePDFScholar
2015

Grounding English Commands to Reward Functions

RSS 2015poster

As intelligent robots become more prevalent, methods to make interaction with the robots more accessible are increasingly important. Communicating the tasks that a person wants the robot to carry out via natural language, and training the robot to ground the natural language through demonstration, a…