← Search

Christina Q Knight

3 accepted papers

2026

Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs

ICLR 2026poster

Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through \textit{elicitati…

Cited by 0SourceScholar
2026

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

ICLR 2026poster

As AI systems progresses, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come to those decisions. Reasoning language models, which provid…

Cited by 0SourcecodeScholar
2026

Reliable Weak-to-Strong Monitoring of LLM Agents

ICLR 2026oral

We stress test monitoring systems for detecting covert misbehavior in LLM agents (e.g., secretly exfiltrating data). We propose a monitor red teaming (MRT) workflow that varies agent and monitor awareness, adversarial evasion strategies, and evaluation across tool-calling (SHADE-Arena) and computer-…

Cited by 0SourcecodeScholar