← Search

Neemesh Yadav

4 accepted papers

2025

Inference-Time Selective Debiasing to Enhance Fairness in Text Classification Models

NAACL 2025short

We propose selective debiasing – an inference-time safety mechanism designed to enhance the overall model quality in terms of prediction performance and fairness, especially in scenarios where retraining the model is impractical. The method draws inspiration from selective classification, where at i…

Cited by 0SourcePDFScholar
2025

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs

COLING 2025main

The rise of large language models (LLMs) has created a need for advanced benchmarking systems beyond traditional setups. To this end, we introduce QUENCH, a novel text-based English Quizzing Benchmark manually curated and transcribed from YouTube quiz videos. QUENCH possesses masked entities and rat…

2025

Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing

ACL 2025finding

The ability of Natural Language Processing (NLP) methods to categorize text into multiple classes has motivated their use in online content moderation tasks, such as hate speech and fake news detection. However, there is limited understanding of how or why these methods make such decisions, or why c…

2024

Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech

ACL 2024findings

Employing language models to generate explanations for an incoming implicit hate post is an active area of research. The explanation is intended to make explicit the underlying stereotype and aid content moderators. The training often combines top-k relevant knowledge graph (KG) tuples to provide wo…