2025
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
NeurIPS 2025poster
Despite the impressive performance of general-purpose large language models (LLMs), they often require fine-tuning or post-training to excel at specific tasks. For instance, large reasoning models (LRMs), such as the DeepSeek-R1 series, demonstrate strong reasoning capabilities after post-train…