← Search

Nils Lukas

12 accepted papers

2026

Addressing Overcommitment in the Reasoning of Gendered Economic Memes Under Multimodal Ambiguity

IJCAI 2026

Multimodal meme understanding is increasingly used to analyze socially sensitive content, yet existing models often exhibit biased behavior when interpreting economic dependence and social roles under ambiguity. Many memes express economic relationships through sparse text or symbolic visual cues, p

Cited by 0Scholar
2026

CoRe: Collaborative Reasoning via Cross Teaching

ICML 2026poster

Large language models exhibit complementary reasoning errors: on the same instance, one model may succeed with a particular decomposition while another fails. We propose Collaborative Reasoning (CORE), a training-time collaboration framework that converts peer success into a learning signal via a cr…

Cited by 0SourceScholar
2026

DP-Fusion: Token-Level Differentially Private Inference for Large Language Models

ICLR 2026poster

Large language models (LLMs) do not preserve privacy at inference-time. The LLM's outputs can inadvertently reveal information about the model's context, which presents a privacy challenge when the LLM is augmented via tools or databases containing sensitive information. Existing privacy-preserving…

Cited by 0SourcecodeScholar
2025

Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems

ICML 2025poster

Vision Language Model (VLM) Agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent systems comprise specialized agents who collaborate to solve a (complex) task. A core security property is **robustness**, stat…

Cited by 0SourcePDFScholar
2025

Optimizing Adaptive Attacks against Watermarks for Language Models

ICML 2025spotlight

Large Language Models (LLMs) can be misused to spread unwanted content at scale. Content watermarking deters misuse by hiding messages in content, enabling its detection using a secret *watermarking key*. Robustness is a core security property, stating that evading detection requires (significant) d…

2025

SPIRIT: Patching Speech Language Models against Jailbreak Attacks

EMNLP 2025

Speech Language Models (SLMs) enable natural interactions via spoken instructions, which more effectively capture user intent by detecting nuances in speech. The richer speech signal introduces new security risks compared to text-based models, as adversaries can better bypass safety mechanisms by in

Cited by 0SourcePDFScholar
2024

Leveraging Optimization for Adaptive Attacks on Image Watermarks

ICLR 2024poster

Untrustworthy users can misuse image generators to synthesize high-quality deepfakes and engage in unethical activities. Watermarking deters misuse by marking generated content with a hidden message, enabling its detection using a secret watermarking key. A core security property of watermarking is…

2021

Deep Neural Network Fingerprinting by Conferrable Adversarial Examples

ICLR 2021spotlight

In Machine Learning as a Service, a provider trains a deep neural network and gives many users access. The hosted (source) model is susceptible to model stealing attacks, where an adversary derives a surrogate model from API access to the source model. For post hoc detection of such attacks, the pro…

Cited by 186SourcePDFScholar