← Search

Artur Zolkowski

1 accepted papers

2026

Early Signs of Steganographic Capabilities in Frontier LLMs

ICLR 2026poster

Monitoring Large Language Model (LLM) outputs is crucial for mitigating risks from misuse and misalignment. However, LLMs could evade monitoring through steganography: Encoding hidden information within seemingly benign generations. In this paper, we evaluate the steganography capabilities in fronti…

Cited by 0SourcecodeScholar