2024
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
EMNLP 2024main
Large Language Models (LLMs) excel in diverse tasks but often underperform in specialized fields due to limited domain-specific or proprietary corpus. Continual pre-training (CPT) enhances LLM capabilities by imbuing new domain-specific or proprietary knowledge while replaying general corpus to prev…