← Search

Wai Man Si

2 accepted papers

2025

Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms

NeurIPS 2025poster

Despite the impressive performance of general-purpose large language models (LLMs), they often require fine-tuning or post-training to excel at specific tasks. For instance, large reasoning models (LRMs), such as the DeepSeek-R1 series, demonstrate strong reasoning capabilities after post-train…

Cited by 0SourceScholar
2025

SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation

ICLR 2025poster

As advancements in large language models (LLMs) continue and the demand for personalized models increases, parameter-efficient fine-tuning (PEFT) methods (e.g., LoRA) become essential due to their efficiency in reducing computation costs. However, recent studies have raised alarming concerns that Lo…

Cited by 3SourcePDFScholar