← Search

Ruichao Mo

2 accepted papers

2026

Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink

ICML 2026spotlight

Harmful fine-tuning can invalidate safety alignment of large language models, exposing significant safety risks. In this paper, we utilize the attention sink mechanism to mitigate harmful fine-tuning. Specifically, we first measure a statistic named *sink divergence* for each attention head and obse…

Cited by 0SourceScholar
2024

CycleNet: Enhancing Time Series Forecasting through Modeling Periodic Patterns

NeurIPS 2024spotlight

The stable periodic patterns present in time series data serve as the foundation for conducting long-horizon forecasts. In this paper, we pioneer the exploration of explicitly modeling this periodicity to enhance the performance of models in long-term time series forecasting (LTSF) tasks. Specifical…