← Search

Firoj Alam

24 accepted papers

2025

Annotating the Annotators: Analysis, Insights and Modelling from an Annotation Campaign on Persuasion Techniques Detection

ACL 2025finding

Persuasion (or propaganda) techniques detection is a relatively novel task in Natural Language Processing (NLP). While there have already been a number of annotation campaigns, they have been based on heuristic guidelines, which have never been thoroughly discussed. Here, we present the first system…

Cited by 4SourcePDFScholar
2025

AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs

COLING 2025main

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern Standard Arabic (MSA), created using Machine Translation (MT) c…

Cited by 7SourcePDFScholar
2025

BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting

NAACL 2025findings

This paper introduces BnTTS (Bangla Text-To-Speech), the first framework for Bangla speaker adaptation-based TTS, designed to bridge the gap in Bangla speech synthesis using minimal training data. Building upon the XTTS architecture, our approach integrates Bangla into a multilingual TTS pipeline, w…

Cited by 0SourcePDFScholar
2025

LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content

NAACL 2025findings

Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. However, their capabilities remain limited when addressing domain-specific problems, particularly in downstream NLP tasks. Research has shown that models fine-tuned on instruction…

Cited by 3SourcePDFScholar
2025

MemeIntel: Explainable Detection of Propagandistic and Hateful Memes

EMNLP 2025

The proliferation of multimodal content on social media presents significant challenges in understanding and moderating complex, context-dependent issues such as misinformation, hate speech, and propaganda. While efforts have been made to develop resources and propose new methods for automatic detec

2025

NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

ACL 2025finding

Natural Question Answering (QA) datasets play a crucial role in evaluating the capabilities of large language models (LLMs), ensuring their effectiveness in real-world applications. Despite the numerous QA datasets that have been developed and some work done in parallel, there is a notable lack of a…

Cited by 0SourcePDFScholar
2025

PropXplain: Can LLMs Enable Explainable Propaganda Detection?

EMNLP 2025

There has been significant research on propagandistic content detection across different modalities and languages. However, most studies have primarily focused on detection, with little attention given to explanations justifying the predicted label. This is largely due to the lack of resources that

2025

TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking

ACL 2025finding

In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes. Due to computational constraints during both training and inference, we focused on smaller models. To train TituLLMs, we collected a pretraining dataset of approximately ∼ 37 billion t…

2024

ArMeme: Propagandistic Content in Arabic Memes

EMNLP 2024main

With the rise of digital communication memes have become a significant medium for cultural and political expression that is often used to mislead audience. Identification of such misleading and persuasive multimodal content become more important among various stakeholders, including social media pla…

2024

Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles

COLING 2024main

The use of propaganda has spiked on mainstream and social media, aiming to manipulate or mislead users. While efforts to automatically detect propaganda techniques in textual, visual, or multimodal content have increased, most of them primarily focus on English content. The majority of the recent in…

Cited by 27SourcePDFScholar
2024

Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis

COLING 2024main

The rapid expansion of the digital world has propelled sentiment analysis into a critical tool across diverse sectors such as marketing, politics, customer service, and healthcare. While there have been significant advancements in sentiment analysis for widely spoken languages, low-resource language…

2023

ConceptX: A Framework for Latent Concept Analysis

AAAI 2023technical

The opacity of deep neural networks remains a challenge in deploying solutions where explanation is as important as precision. We present ConceptX, a human-in-the-loop framework for interpreting and annotating latent representational space in pre-trained Language Models (pLMs). We use an unsupervise…

2022

A Survey on Multimodal Disinformation Detection

COLING 2022main

Recent years have witnessed the proliferation of offensive content online such as fake news, propaganda, misinformation, and disinformation. While initially this was mostly about textual content, over time images and videos gained popularity, as they are much easier to consume, attract more attentio…

Cited by 156SourcePDFScholar
2022

Analyzing Encoded Concepts in Transformer Language Models

NAACL 2022long

We propose a novel framework ConceptX, to analyze how latent concepts are encoded in representations learned within pre-trained lan-guage models. It uses clustering to discover the encoded concepts and explains them by aligning with a large set of human-defined concepts. Our analysis on seven transf…

2022

Assisting the Human Fact-Checkers: Detecting All Previously Fact-Checked Claims in a Document

EMNLP 2022finding

Given the recent proliferation of false claims online, there has been a lot of manual fact-checking effort. As this is very time-consuming, human fact-checkers can benefit from tools that can support them and make them more efficient. Here, we focus on building a system that could provide such suppo…

2022

Detecting and Understanding Harmful Memes: A Survey

IJCAI 2022poster

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content often combines multiple modalities, as in the case of memes.…

2022

Discovering Latent Concepts Learned in BERT

ICLR 2022poster

A large number of studies that analyze deep neural network models and their ability to encode various linguistic and non-linguistic concepts provide an interpretation of the inner mechanics of these models. The scope of the analyses is limited to pre-defined concepts that reinforce the traditional l…

Cited by 77SourcePDFScholar
2022

Effect of Post-processing on Contextualized Word Representations

COLING 2022main

Post-processing of static embedding has been shown to improve their performance on both lexical and sequence-level tasks. However, post-processing for contextualized embeddings is an under-studied problem. In this work, we question the usefulness of post-processing for contextualized embeddings obta…

Cited by 17SourcePDFScholar
2022

On the Transformation of Latent Space in Fine-Tuned NLP Models

EMNLP 2022main

We study the evolution of latent space in fine-tuned NLP models. Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations. More specifically, we discover latent concepts in the representational space using hierarchical clustering. We then use a…

Cited by 23SourcePDFScholar
2022

The Role of Context in Detecting Previously Fact-Checked Claims

NAACL 2022findings

Recent years have seen the proliferation of disinformation and fake news online. Traditional approaches to mitigate these issues is to use manual or automatic fact-checking. Recently, another approach has emerged: checking whether the input claim has previously been fact-checked, which can be done a…

2021

Automated Fact-Checking for Assisting Human Fact-Checkers

IJCAI 2021poster

The reporting and the analysis of current events around the globe has expanded from professional, editor-lead journalism all the way to citizen journalism. Nowadays, politicians and other key players enjoy direct access to their audiences through social media, bypassing the filters of official cable…

Cited by 281SourcePDFScholar
2021

Detecting Propaganda Techniques in Memes

ACL 2021long

Propaganda can be defined as a form of communication that aims to influence the opinions or the actions of people towards a specific goal; this is achieved by means of well-defined rhetorical and psychological devices. Propaganda, in the form we know it today, can be dated back to the beginning of t…

2021

Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society

EMNLP 2021finding

With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic. Fighting this infodemic has been declared one of the most important focus areas of the World Health Organ…