Dissecting Post-Training: Uncovering the Complementary Roles of SFT and RL for Document Parsing
Document parsing, the task of extracting diverse content from PDFs while preserving their structural integrity, has been significantly advanced by Multimodal Large Language Models (MLLMs). These models have achieved remarkable success, largely driven by extensive post-training on massive datasets. T…