ICLR 2026poster0 citations

CIMemories: A Compositional Benchmark For Contextual Integrity In LLMs

Niloofar Mireshghallah, Neal Mangaokar, Narine Kokhlikyan, Arman Zharmagambetov, Manzil Zaheer, Saeed Mahloujifar, Kamalika Chaudhuri

Abstract

Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory creates critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context. CIMemories uses synthetic user profiles with 100+ attributes per user, paired with various task contexts where each attribute may be essential for some tasks but inappropriate for others. For example, mental health details are necessary for booking therapy but inappropriate when requesting time off from work. This design enables two forms of compositionality: (1) flexible memory composition by varying which attributes are necessary versus inappropriate across different settings, and (2) multi-task composition per user, measuring cumulative information disclosure across sessions. Our evaluation reveals frontier models exhibit between 14%-69% attribute-level violations (leaking inappropriate information), and that higher task completeness (sharing necessary information) is accompanied by increased violations, highlighting critical gaps in integrity-aware memory systems.

Contextual IntegrityInference-time PrivacyInput-output flow
BibTeX
@inproceedings{
mireshghallah2026cimemories,
title={{CIM}emories: A Compositional Benchmark For Contextual Integrity In {LLM}s},
author={Niloofar Mireshghallah and Neal Mangaokar and Narine Kokhlikyan and Arman Zharmagambetov and Manzil Zaheer and Saeed Mahloujifar and Kamalika Chaudhuri},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=YnNIp38v1M}
}
CIMemories: A Compositional Benchmark For Contextual Integrity In LLMs · ICLR 2026