← Search

Palash Nandi

2 accepted papers

2025

SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection

EMNLP 2025

Large Language Models (LLMs) with safe-alignment training are powerful instruments with robust language comprehension capability. Typically LLMs undergo careful alignment training involving human feedback to ensure the acceptance of safe inputs while rejection of harmful or unsafe ones. However, the

2024

Recent Advances in Online Hate Speech Moderation: Multimodality and the Role of Large Models

EMNLP 2024finding

Moderating hate speech (HS) in the evolving online landscape is a complex challenge, compounded by the multimodal nature of digital content. This survey examines recent advancements in HS moderation, focusing on the burgeoning role of large language models (LLMs) and large multimodal models (LMMs) i…

Cited by 1SourcePDFScholar