← Search

Filip Sondej

2 accepted papers

2025

How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

EMNLP 2025

Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations—attributing its effects solely to dampened toxic neurons in the MLP layers—are incomplete

Cited by 0SourcePDFScholar
2025

Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems

AAAI 2025technical

As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agents collaborating on shared objectives to study these security risks and security trade-offs. We focus on scenarios where…

Cited by 1SourcePDFScholar