← Search

Jakub Podolak

2 accepted papers

2025

Gandalf the Red: Adaptive Security for LLMs

ICML 2025poster

Current evaluations of defenses against prompt attacks in large language model (LLM) applications often overlook two critical factors: the dynamic nature of adversarial behavior and the usability penalties imposed on legitimate users by restrictive defenses. We propose D-SEC (Dynamic Security Utilit…

2024

LLM generated responses to mitigate the impact of hate speech

EMNLP 2024finding

In this study, we explore the use of Large Language Models (LLMs) to counteract hate speech. We conducted the first real-life A/B test assessing the effectiveness of LLM-generated counter-speech. During the experiment, we posted 753 automatically generated responses aimed at reducing user engagement…

Cited by 3SourcePDFScholar