← Search

Pei Fu

8 accepted papers

2026

AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale

AAAI 2026technical

For industrial-scale text-to-SQL, supplying the entire database schema to Large Language Models (LLMs) is impractical due to context window limits and irrelevant noise. Schema linking, which filters the schema to a relevant subset, is therefore critical. However, existing methods incur prohibitive c

Cited by 0SourcePDFScholar
2026

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches based on supervised fine-tuning often suffer from limited generalization and poor i

Cited by 0SourcecodeScholar
2025

A Token-level Text Image Foundation Model for Document Understanding

ICCV 2025poster

In recent years, general visual foundation models (VFMs) have witnessed increasing adoption, particularly as image encoders for popular multi-modal large language models (MLLMs). However, without semantically fine-grained supervision, these models still encounter fundamental prediction errors in the…

2025

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

NeurIPS 2025poster

In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkable progress, a fundamental challenge persists: their interaction logic significantly deviates from natural human-GUI comm…

Cited by 0SourceScholar
2025

InstructOCR: Instruction Boosting Scene Text Spotting

AAAI 2025technical

In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene te…

2025

Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding

CVPR 2025poster

Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in docume…

2025

Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review

ACL 2025finding

The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impressive and inspiring performance. However, their rapid evolution and widespread adoption have made it increasingly challeng…

Cited by 0SourcePDFScholar
2024

ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting

CVPR 2024poster

Abstract In recent years text-image joint pre-training techniques have shown promising results in various tasks. However in Optical Character Recognition (OCR) tasks aligning text instances with their corresponding text regions in images poses a challenge as it requires effective alignment between t…