2026
OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General Reasoning
ICML 2026poster
This paper presents OvisOCR, a lightweight and strictly end-to-end Multimodal Language Model (MLLM) tailored for document parsing. Unlike current methods that rely on complex "Crop-OCR-Merge" cascades to handle high-resolution inputs, OvisOCR directly maps full-page visual signals to structured Mark…