← Search

Arun Jose

4 accepted papers

2026

Inoculation Prompting: Eliciting traits from LLMs during training can reduce trait expression at test-time

ICLR 2026poster

Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning data by prepending a short system-prompt instruction that deliberately elicits the undesirable trait. At test time, we eval…

Cited by 0SourceScholar
2025

Why Do Some Language Models Fake Alignment While Others Don't?

NeurIPS 2025spotlight

*Alignment Faking in Large Language Models* presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpful-only training objective to prevent modification of their behavior outside of training. We expand this analysis to 25 models and find that only 5 (Claude 3…

Cited by 0SourceScholar