← Search

David Glukhov

2 accepted papers

2025

Breach By A Thousand Leaks: Unsafe Information Leakage in 'Safe' AI Responses

ICLR 2025poster

Vulnerability of Frontier language models to misuse has prompted the development of safety measures like filters and alignment training seeking to ensure safety through robustness to adversarially crafted prompts. We assert that robustness is fundamentally insufficient for ensuring safety goals due…

Cited by 2SourcePDFScholar
2024

Position: Fundamental Limitations of LLM Censorship Necessitate New Approaches

ICML 2024poster

Large language models (LLMs) have exhibited impressive capabilities in comprehending complex instructions. However, their blind adherence to provided instructions has led to concerns regarding risks of malicious use. Existing defence mechanisms, such as model fine-tuning or output censorship methods…

Cited by 2SourcePDFScholar