← Search

Feiyu Gao

10 accepted papers

2025

A Simple yet Effective Layout Token in Large Language Models for Document Understanding

CVPR 2025poster

Recent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a metho…

Cited by 0SourcePDFScholar
2025

Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Grounding

EMNLP 2025

In the daily work, vast amounts of documents are stored in pixel-based formats such as images and scanned PDFs, posing challenges for efficient database management and data processing. Existing methods often fragment the parsing process into the pipeline of separated subtasks on the layout element l

Cited by 0SourcePDFScholar
2025

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

EMNLP 2025

Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand. As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities. However, due to

Cited by 0SourcePDFScholar
2024

DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing

EMNLP 2024main

Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding. However, previously the research on this topic has been largely hindered since most existing datasets are small-s…

2024

Visual Text Generation in the Wild

ECCV 2024poster

"Recently, with the rapid advancements of generative models, the field of visual text generation has witnessed significant progress. However, it is still challenging to render high-quality text images in real-world scenarios, as three critical criteria should be satisfied: (1) Fidelity: the generate…

2024

WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation

ECCV 2024poster

"In the era of content creation revolution propelled by advancements in generative models, the field of web design remains unexplored despite its critical role in modern digital communication. The web design process is complex and often time-consuming, especially for those with limited expertise. In…

2023

GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree

EMNLP 2023long main

Inexhaustible web content carries abundant perceptible information beyond text. Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyber-richness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style…

Cited by 0SourceScholar
2023

LORE: Logical Location Regression Network for Table Structure Recognition

AAAI 2023technical

Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they e…

2020

An End-to-End OCR Text Re-organization Sequence Learning for Rich-text Detail Image Comprehension

ECCV 2020poster

Nowadays rich description on detail images help users know more about the commodities. With the help of OCR technology, the description text can be detected and recognized as auxiliary information to remove the comprehending barriers among the visual impaired users. However, for lack of proper logic…

Cited by 29SourcePDFScholar