← Search

Songxuan Lai

2 accepted papers

2026

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

ICLR 2026poster

Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. However, their capabilities in text-rich image reasoning tasks remain understudied due to the absence of a dedicated and systematic benchmark. To address this gap,…

Cited by 0SourcecodeScholar
2025

DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

AAAI 2025technical

Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layouts typical of document images. These characteristics demand a high level of detail perception ability from MLLMs. While i…

Cited by 8SourcePDFScholar