← Search

Soroush Vosoughi

62 accepted papers

2026

Diffusion Language Model Knows the Answer Before It Decodes

ICLR 2026oral

Diffusion language models (DLMs) have recently emerged as an alternative to autoregressive approaches, offering parallel sequence generation and flexible token orders. However, their inference remains slower than that of autoregressive models, primarily due to the cost of bidirectional attention and…

Cited by 0SourcecodeScholar
2025

A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling

EMNLP 2025

Rhetorical strategies are central to persuasive communication, from political discourse and marketing to legal argumentation. However, analysis of rhetorical strategies has been limited by reliance on human annotation, which is costly, inconsistent, difficult to scale. Their associated datasets are

Cited by 0SourcePDFScholar
2025

Communication Makes Perfect: Persuasion Dataset Construction via Multi-LLM Communication

NAACL 2025long

Large Language Models (LLMs) have shown proficiency in generating persuasive dialogue, yet concerns about the fluency and sophistication of their outputs persist. This paper presents a multi-LLM communication framework designed to enhance the generation of persuasive data automatically. This framewo…

Cited by 1SourcePDFScholar
2025

DECASTE: Unveiling Caste Stereotypes in Large Language Models Through Multi-Dimensional Bias Analysis

IJCAI 2025

Recent advancements in large language models (LLMs) have revolutionized natural language processing (NLP) and expanded their applications across diverse domains. However, despite their impressive capabilities, LLMs have been shown to reflect and perpetuate harmful societal biases, including those ba

2025

Enhancing LLM-Based Persuasion Simulations with Cultural and Speaker-Specific Information

EMNLP 2025

Large language models (LLMs) have been used to synthesize persuasive dialogues for studying persuasive behavior. However, existing approaches often suffer from issues such as stance oscillation and low informativeness. To address these challenges, we propose reinforced instructional prompting, a met

Cited by 0SourcePDFScholar
2025

Growing Through Experience: Scaling Episodic Grounding in Language Models

ACL 2025long

Language models (LMs) require effective episodic grounding—the ability to learn from and apply past experiences—to perform well at physical planning tasks. While current approaches struggle with scalability and integration of episodic memory, which is particularly limited for medium-sized LMs (7B pa…

Cited by 0SourcePDFScholar
2025

ImpScore: A Learnable Metric For Quantifying The Implicitness Level of Sentences

ICLR 2025spotlight

Handling implicit language is essential for natural language processing systems to achieve precise text understanding and facilitate natural interactions with users. Despite its importance, the absence of a metric for accurately measuring the implicitness of language significantly constrains the dep…

2025

Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages

NAACL 2025short

Endangered languages, such as Navajo—the most widely spoken Native American language—are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. This study evaluates Google’s Language Identification (LangID) tool, wh…

Cited by 2SourcePDFScholar
2025

Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judge

EMNLP 2025

LLM-as-Judge has emerged as a scalable alternative to human evaluation, enabling large language models (LLMs) to provide reward signals in trainings. While recent work has explored multi-agent extensions such as multi-agent debate and meta-judging to enhance evaluation quality, the question of how i

Cited by 0SourcePDFScholar
2025

Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making

EMNLP 2025

Modern embodied AI uses multimodal large language models (MLLMs) as policy models, predicting actions from final-layer hidden states. This widely adopted approach, however, assumes that monolithic last-layer representations suffice for decision-making—a structural simplification at odds with decades

Cited by 0SourcePDFScholar
2025

Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner

ICML 2025spotlight

Theory-of-mind (ToM) enables humans to infer mental states—such as beliefs, desires, and intentions—forming the foundation of social cognition. Existing computational ToM methods rely on structured workflows with ToM-specific priors or deep model fine-tuning but struggle with scalability in multimod…

Cited by 0SourcePDFScholar
2025

Pretrained Image-Text Models are Secretly Video Captioners

NAACL 2025short

Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these sequences. However, we find that by using minimal computational resources and without complex modifications to address vide…

2025

ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering

EMNLP 2025

Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous systems, where models must provide not only accurate answers but also explanations that humans can easily understand an

Cited by 0SourcePDFScholar
2025

Recontextualizing Revitalization: A Mixed Media Approach to Reviving the Nüshu Language

EMNLP 2025

Nüshu is an endangered language from Jiangyong County, China, and the world’s only known writing system created and used exclusively by women. Recent Natural Language Processing (NLP) work has digitized small Nüshu-Chinese corpora, but the script remains computationally inaccessible due to its handw

Cited by 0SourcePDFScholar
2025

Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

EMNLP 2025

Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures. To address this gap, we introduce a cost-efficient human-LLM collaborative ann

Cited by 0SourcePDFScholar
2025

SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models

EMNLP 2025

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic approach that involves a capable base model, high-quality rea

2025

Superficial Self-Improved Reasoners Benefit from Model Merging

EMNLP 2025

Large Language Models (LLMs) rely heavily on large-scale reasoning data, but as such data becomes increasingly scarce, model self-improvement offers a promising alternative. However, this process can lead to model collapse, as the model’s output becomes overly deterministic with reduced diversity. I

2025

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

NAACL 2025findings

Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models face inherent limitations due to their finite internal capacity, which restricts their ability to process extended tempora…

2025

Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification

ACL 2025finding

Indigenous languages remain largely invisible in commercial language identification (LID) systems, a stark reality exemplified by Google Translate’s LangID tool, which supports over 100 languages but excludes all 150 Indigenous languages of North America. This technological marginalization is partic…

Cited by 0SourcePDFScholar
2024

Achieving Domain-Independent Certified Robustness via Knowledge Continuity

NeurIPS 2024poster

We present *knowledge continuity*, a novel definition inspired by Lipschitz continuity which aims to certify the robustness of neural networks across input domains (such as continuous and discrete domains in vision and language, respectively). Most existing approaches that seek to certify robustness…

2024

Addressing Healthcare-related Racial and LGBTQ+ Biases in Pretrained Language Models

NAACL 2024findings

Recent studies have highlighted the issue of Pretrained Language Models (PLMs) inadvertently propagating social stigmas and stereotypes, a critical concern given their widespread use. This is particularly problematic in sensitive areas like healthcare, where such biases could lead to detrimental out…

Cited by 3SourcePDFScholar
2024

AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality

EMNLP 2024main

Parameter-efficient fine-tuning methods, such as Low-Rank Adaptation (LoRA), are known to enhance training efficiency in Large Language Models (LLMs). Due to the limited parameters of LoRA, recent studies seek to combine LoRA with Mixture-of-Experts (MoE) to boost performance across various tasks. H…

2024

Disordered-DABS: A Benchmark for Dynamic Aspect-Based Summarization in Disordered Texts

EMNLP 2024finding

Aspect-based summarization has seen significant advancements, especially in structured text. Yet, summarizing disordered, large-scale texts, like those found in social media and customer feedback, remains a significant challenge. Current research largely targets predefined aspects within structured…

2024

Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction

ACL 2024long

We introduce EVLGen, a streamlined framework designed for the pre-training of visually conditioned language generation models with high computational demands, utilizing frozen pre-trained large language models (LLMs). The conventional approach in vision-language pre-training (VLP) typically involves…

2024

GEM: Generating Engaging Multimodal Content

IJCAI 2024poster

Generating engaging multimodal content is a key objective in numerous applications, such as the creation of online advertisements that captivate user attention through a synergy of images and text. In this paper, we introduce GEM, a novel framework engineered for the generation of engaging multimoda…

Cited by 1SourcePDFScholar
2024

Interpretable Image Classification with Adaptive Prototype-based Vision Transformers

NeurIPS 2024poster

We present ProtoViT, a method for interpretable image classification combining deep learning and case-based reasoning. This method classifies an image by comparing it to a set of learned prototypes, providing explanations of the form ``this looks like that.'' In our model, a prototype consists of **…

2024

Is GPT-4V (ision) All You Need for Automating Academic Data Visualization? Exploring Vision-Language Models’ Capability in Reproducing Academic Charts

EMNLP 2024finding

While effective data visualization is crucial to present complex information in academic research, its creation demands significant expertise in both data management and graphic design. We explore the potential of using Vision-Language Models (VLMs) in automating the creation of data visualizations…

Cited by 2SourcePDFScholar
2024

MentalManip: A Dataset For Fine-grained Analysis of Mental Manipulation in Conversations

ACL 2024long

Mental manipulation, a significant form of abuse in interpersonal conversations, presents a challenge to identify due to its context-dependent and often subtle nature. The detection of manipulative language is essential for protecting potential victims, yet the field of Natural Language Processing (…

2024

Simulated Misinformation Susceptibility (SMISTS): Enhancing Misinformation Research with Large Language Model Simulations

ACL 2024findings

Psychological inoculation, a strategy designed to build resistance against persuasive misinformation, has shown efficacy in curbing its spread and mitigating its adverse effects at early stages. Despite its effectiveness, the design and optimization of these inoculations typically demand substantial…

Cited by 2SourcePDFScholar
2024

The Computational Anatomy of Humility: Modeling Intellectual Humility in Online Public Discourse

EMNLP 2024main

The ability for individuals to constructively engage with one another across lines of difference is a critical feature of a healthy pluralistic society. This is also true in online discussion spaces like social media platforms. To date, much social media research has focused on preventing ills—like…

2024

Training Socially Aligned Language Models on Simulated Social Interactions

ICLR 2024poster

The goal of social alignment for AI systems is to make sure these models can conduct themselves appropriately following social values. Unlike humans who establish a consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly recite the corpus in soci…

2024

Working Memory Identifies Reasoning Limits in Language Models

EMNLP 2024main

This study explores the inherent limitations of large language models (LLMs) from a scaling perspective, focusing on the upper bounds of their cognitive capabilities. We integrate insights from cognitive science to quantitatively examine how LLMs perform on n-back tasks—a benchmark used to assess wo…

Cited by 8SourcePDFScholar
2023

Bootstrapping Vision-Language Learning with Decoupled Language Pre-training

NeurIPS 2023spotlight

We present a novel methodology aimed at optimizing the application of frozen large language models (LLMs) for resource-intensive vision-language (VL) pre-training. The current paradigm uses visual features as prompts to guide language models, with a focus on determining the most relevant visual feat…

2023

Deciphering Stereotypes in Pre-Trained Language Models

EMNLP 2023long main

Warning: This paper contains content that is stereotypical and may be upsetting. This paper addresses the issue of demographic stereotypes present in Transformer-based pre-trained language models (PLMs) and aims to deepen our understanding of how these biases are encoded in these models. To accompl…

Cited by 0SourceScholar
2023

Improving Representation Learning for Histopathologic Images with Cluster Constraints

ICCV 2023poster

Recent advances in whole-slide image (WSI) scanners and computational capabilities have significantly propelled the application of artificial intelligence in histopathology slide analysis. While these strides are promising, current supervised learning approaches for WSI analysis come with the challe…

Cited by 18PDFcodeScholar
2023

Improving Syntactic Probing Correctness and Robustness with Control Tasks

ACL 2023short

Syntactic probing methods have been used to examine whether and how pre-trained language models (PLMs) encode syntactic features. However, the probing methods are usually biased by the PLMs’ memorization of common word co-occurrences, even if they do not form syntactic relations. This paper presents…

Cited by 2SourcePDFScholar
2023

Intersectional Stereotypes in Large Language Models: Dataset and Analysis

EMNLP 2023short findings

Despite many stereotypes targeting intersectional demographic groups, prior studies on stereotypes within Large Language Models (LLMs) primarily focus on broader, individual categories. This research bridges this gap by introducing a novel dataset of intersectional stereotypes, curated with the assi…

Cited by 0SourceScholar
2023

Language models are multilingual chain-of-thought reasoners

ICLR 2023poster

We evaluate the reasoning abilities of large language models in multilingual settings. We introduce the Multilingual Grade School Math (MGSM) benchmark, by manually translating 250 grade-school math problems from the GSM8K dataset (Cobbe et al., 2021) into ten typologically diverse languages. We fin…

2023

Mind's Eye: Grounded Language Model Reasoning through Simulation

ICLR 2023poster

Successful and effective communication between humans and AI relies on a shared experience of the world. By training solely on written text, current language models (LMs) miss the grounded experience of humans in the real-world---their failure to relate language to the physical world causes knowledg…

Cited by 84SourcePDFScholar
2023

Proto-lm: A Prototypical Network-Based Framework for Built-in Interpretability in Large Language Models

EMNLP 2023long findings

Large Language Models (LLMs) have significantly advanced the field of Natural Language Processing (NLP), but their lack of interpretability has been a major concern. Current methods for interpreting LLMs are post hoc, applied after inference time, and have limitations such as their focus on low-leve…

Cited by 0SourcecodeScholar
2022

Aligning Generative Language Models with Human Values

NAACL 2022findings

Although current large-scale generative language models (LMs) can show impressive insights about factual knowledge, they do not exhibit similar success with respect to human values judgements (e.g., whether or not the generations of an LM are moral). Existing methods learn human values either by dir…

2022

Contrastive Learning for Prompt-based Few-shot Language Learners

NAACL 2022long

The impressive performance of GPT-3 using natural language prompts and in-context learning has inspired work on better fine-tuning of moderately-sized models under this paradigm. Following this line of work, we present a contrastive learning framework that clusters inputs from the same class for bet…

2022

EnCBP: A New Benchmark Dataset for Finer-Grained Cultural Background Prediction in English

ACL 2022findings

While cultural backgrounds have been shown to affect linguistic expressions, existing natural language processing (NLP) research on culture modeling is overly coarse-grained and does not examine cultural differences among speakers of the same language. To address this problem and augment NLP models…

Cited by 7SourcePDFScholar
2022

Knowledge Infused Decoding

ICLR 2022poster

Pre-trained language models (LMs) have been shown to memorize a substantial amount of knowledge from the pre-training corpora; however, they are still limited in recalling factually correct knowledge given a certain context. Hence. they tend to suffer from counterfactual or hallucinatory generation…

2022

Non-Linguistic Supervision for Contrastive Learning of Sentence Embeddings

NeurIPS 2022accept

Semantic representation learning for sentences is an important and well-studied problem in NLP. The current trend for this task involves training a Transformer-based sentence encoder through a contrastive objective with text, i.e., clustering sentences with semantically similar meanings and scatteri…

2022

Non-Parallel Text Style Transfer with Self-Parallel Supervision

ICLR 2022poster

The performance of existing text style transfer models is severely limited by the non-parallel datasets on which the models are trained. In non-parallel datasets, no direct mapping exists between sentences of the source and target style; the style transfer models thus only receive weak supervision o…

2022

Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits

NeurIPS 2022accept

We present Second Thoughts, a new learning paradigm that enables language models (LMs) to re-align with human values. By modeling the chain-of-edits between value-unaligned and value-aligned text, with LM fine-tuning and additional refinement through reinforcement learning, Second Thoughts not only…

Cited by 38SourcePDFScholar
2022

TWEETSPIN: Fine-grained Propaganda Detection in Social Media Using Multi-View Representations

NAACL 2022long

Recently, several studies on propaganda detection have involved document and fragment-level analyses of news articles. However, there are significant data and modeling challenges dealing with fine-grained detection of propaganda on social media. In this work, we present TWEETSPIN, a dataset containi…

Cited by 36SourcePDFScholar
2021

Contributions of Transformer Attention Heads in Multi- and Cross-lingual Tasks

ACL 2021long

This paper studies the relative importance of attention heads in Transformer-based models to aid their interpretability in cross-lingual and multi-lingual tasks. Prior research has found that only a few attention heads are important in each mono-lingual Natural Language Processing (NLP) task and pru…

2021

Embedding Heterogeneous Networks into Hyperbolic Space Without Meta-path

AAAI 2021technical

Networks found in the real-world are numerous and varied. A common type of network is the heterogeneous network, where the nodes (and edges) can be of different types. Accordingly, there have been efforts at learning representations of these heterogeneous networks in low-dimensional space. However,…

Cited by 30SourcePDFScholar
2021

Few-Shot Text Classification with Triplet Networks, Data Augmentation, and Curriculum Learning

NAACL 2021long

Few-shot text classification is a fundamental NLP task in which a model aims to classify text into a large number of categories, given only a few training examples per category. This paper explores data augmentation—a technique particularly suitable for training with limited data—for this few-shot,…

2021

GradTS: A Gradient-Based Automatic Auxiliary Task Selection Method Based on Transformer Networks

EMNLP 2021main

A key problem in multi-task learning (MTL) research is how to select high-quality auxiliary tasks automatically. This paper presents GradTS, an automatic auxiliary task selection method based on gradient calculation in Transformer-based models. Compared to AUTOSEM, a strong baseline method, GradTS i…

Cited by 8SourcePDFScholar
2021

Linguistic Complexity Loss in Text-Based Therapy

NAACL 2021long

The complexity loss paradox, which posits that individuals suffering from disease exhibit surprisingly predictable behavioral dynamics, has been observed in a variety of both human and animal physiological systems. The recent advent of online text-based therapy presents a new opportunity to analyze…

Cited by 9SourcePDFScholar