← Search

Jianfeng Si

1 accepted papers

2026

Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training

AAAI 2026technical

Current methods for content safety in Large Language Models (LLMs), such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often rely on multi-stage training pipelines and lack fine-grained, post-deployment controllability. To address these limitations, we propos

Cited by 0SourcePDFScholar