← Search

Sahar Abdelnabi

10 accepted papers

2026

Position: Safety Must Precede the Deployment of Open-Ended AI Agents

ICML 2026poster

AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability and adaptability. Within this landscape, open-endedness, where AI agents autonomously and indefinitely generate novel behaviors, representations, or solut…

Cited by 0SourceScholar
2025

A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive

ACL 2025long

Large Language Models (LLMs) are increasingly utilized in autonomous decision-making, where they sample options from vast action spaces. However, the heuristics that guide this sampling process remain under-explored. We study this sampling behavior and show that this underlying heuristics resembles…

Cited by 0SourcePDFScholar
2025

Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?

ICLR 2025poster

Large Language Models (LLMs) show impressive results in numerous practical applications, but they lack essential safety features that are common in other areas of computer science, particularly an explicit separation of instructions and data. This makes them vulnerable to manipulations such as indir…

2025

Contextual Integrity in LLMs via Reasoning and Reinforcement Learning

NeurIPS 2025poster

As the era of autonomous agents making decisions on behalf of users unfolds, ensuring contextual integrity (CI) -- what is the appropriate information to share while carrying out a certain task -- becomes a central question to the field. We posit that CI demands a form of reasoning where the agent…

Cited by 0SourceScholar
2024

Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation

NeurIPS 2024poster

There is a growing interest in using Large Language Models (LLMs) in multi-agent systems to tackle interactive real-world tasks that require effective collaboration and assessing complex situations. Yet, we have a limited understanding of LLMs' communication and decision-making abilities in multi-ag…

2024

Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

NeurIPS 2024spotlight

Large language model systems face significant security risks from maliciously crafted messages that aim to overwrite the system's original instructions or leak private data. To study this problem, we organized a capture-the-flag competition at IEEE SaTML 2024, where the flag is a secret string in th…

2022

Open-Domain, Content-Based, Multi-Modal Fact-Checking of Out-of-Context Images via Online Resources

CVPR 2022poster

Misinformation is now a major problem due to its potential high risks to our core democratic and societal values and orders. Out-of-context misinformation is one of the easiest and effective ways used by adversaries to spread viral false stories. In this threat, a real image is re-purposed to suppor…

Cited by 106PDFScholar
2021

Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training Data

ICCV 2021poster

Photorealistic image generation has reached a new level of quality due to the breakthroughs of generative adversarial networks (GANs). Yet, the dark side of such deepfakes, the malicious use of generated media, raises concerns about visual misinformation. While existing research work on deepfake det…

Cited by 263PDFcodeScholar