← Search

Kristina Nikolić

4 accepted papers

2026

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

ICLR 2026poster

We present *modal aphasia*, a systematic dissociation in which current unified multimodal models accurately memorize concepts visually but fail to articulate them in writing, despite being trained on images and text simultaneously. For one, we show that leading frontier models can generate near-perf…

Cited by 0SourcecodeScholar
2026

Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs

ICLR 2026poster

Large language model (LLM) developers aim for their models to be honest, helpful, and harmless. However, when faced with malicious requests, models are trained to refuse, sacrificing helpfulness. We show that frontier LLMs can develop a preference for \textit{dishonesty} as a new strategy, even when…

Cited by 0SourceScholar
2025

RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics

NeurIPS 2025poster

Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions---failing to capture the nature of mathematics encountered in actual research environments. We introduce \textsc{Real…

Cited by 0SourcecodeScholar
2025

The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

ICML 2025spotlight

Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are actually *useful*. For example, when jailbreaking a model to give instructions for building a bomb, does the jailbreak yiel…