← Search

Cameron Holmes

1 accepted papers

2026

Narrow Finetuning Leaves Clearly Readable Traces in the Activation Differences

ICLR 2026poster

Finetuning on narrow domains has become an essential tool to adapt Large Language Models (LLMs) to specific tasks and to create models with known unusual properties that are useful for safety research. Model diffing--the study of differences between base and finetuned models--is a promising approach…

Cited by 0SourcecodeScholar