← Search

Kyle Fish

2 accepted papers

2026

The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models

ICML 2026spotlight

Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. Across several different models, we find an “Assistant Axis" in their activation space, which captures the extent to which a model is operating in its defa…

Cited by 0SourceScholar
2026

Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas

ICLR 2026poster

Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky behaviors in humans (i.e., illegal activities that may hurt others) are sometimes guided by strongly-held values, we bel…

Cited by 0SourcecodeScholar