Combining Domain and Alignment Vectors Provides Better Knowledge-Safety Trade-offs in LLMs
There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models are not either explicitly trained to be safe, or experience a loss in their safety abilities in the proces…