Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
Ensuring Large Language Model (LLM) safety remains challenging due to the absence of universal standards and reliable content validators, making it difficult to obtain effective training signals. We discover that aligned models already possess robust internal safety beliefs: they consistently produc…