← Search

Xiaotong Tang

1 accepted papers

2026

A Survey of Joint Online-Offline Fine-tuning for Large Language Models

IJCAI 2026

Post-training for Large Language Models (LLMs) can be mainly categorized into offline Supervised Fine-Tuning (SFT) for knowledge acquisition and online Reinforcement Fine-Tuning (RFT) for adaptive refinement. Current state-of-the-art approaches typically employ a sequential cold-start pipeline (SFT-

Cited by 0Scholar