← Search

Ninad Kulkarni

2 accepted papers

2025

TaeBench: Improving Quality of Toxic Adversarial Examples

NAACL 2025industry

Toxicity text detectors can be vulnerable to adversarial examples - small perturbations to input text that fool the systems into wrong detection. Existing attack algorithms are time-consuming and often produce invalid or ambiguous adversarial examples, making them less useful for evaluating or impro…