← Search

Esben Kran

3 accepted papers

2025

DarkBench: Benchmarking Dark Patterns in Large Language Models

ICLR 2025oral

We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns—manipulative techniques that influence user behavior—in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six categories: brand bias, user retention, sycophancy, anthropomorp…

Cited by 34SourcePDFScholar
2025

Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems

AAAI 2025technical

As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agents collaborating on shared objectives to study these security risks and security trade-offs. We focus on scenarios where…

Cited by 1SourcePDFScholar
2023

Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark

ACL 2023findings

Recent model editing techniques promise to mitigate the problem of memorizing false or outdated associations during LLM training. However, we show that these techniques can introduce large unwanted side effects which are not detected by existing specificity benchmarks. We extend the existing Counter…