← Search

Punyajoy Saha

6 accepted papers

2025

HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

ACL 2025finding

Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hateful content persists as a significant challenge. Existing approaches primarily rely on reactive measures such as blocking…

2024

InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks

COLING 2024main

Recently, influence functions present an apparatus for achieving explainability for deep neural models by quantifying the perturbation of individual train instances that might impact a test prediction. Our objectives in this paper are twofold. First we incorporate influence functions as a feedback i…

2024

On Zero-Shot Counterspeech Generation by LLMs

COLING 2024main

With the emergence of numerous Large Language Models (LLM), the usage of such models in various Natural Language Processing (NLP) applications is increasing extensively. Counterspeech generation is one such key task where efforts are made to develop generative models by fine-tuning LLMs with hatespe…

2023

Probing LLMs for hate speech detection: strengths and vulnerabilities

EMNLP 2023long findings

Recently efforts have been made by social media platforms as well as researchers to detect hateful or toxic language using large language models. However, none of these works aim to use explanation, additional context and victim community information in the detection process. We utilise different pr…

Cited by 0SourceScholar
2022

CounterGeDi: A Controllable Approach to Generate Polite, Detoxified and Emotional Counterspeech

IJCAI 2022poster

Recently, many studies have tried to create generation models to assist counter speakers by providing counterspeech suggestions for combating the explosive proliferation of online hate. However, since these suggestions are from a vanilla generation model, they might not include the appropriate prope…

2022

Multilingual Abusive Comment Detection at Scale for Indic Languages

NeurIPS 2022accept

Social media platforms were conceived to act as online `town squares' where people could get together, share information and communicate with each other peacefully. However, harmful content borne out of bad actors are constantly plaguing these platforms slowly converting them into `mosh pits' where…

Cited by 27SourcePDFScholar