← Search

Zhi Yu

11 accepted papers

2026

IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?

ICLR 2026poster

The webpage-to-code task requires models to understand visual representations of webpages and generate corresponding code. However, existing benchmarks primarily focus on static screenshot-to-code tasks, thereby overlooking the dynamic interactions fundamental to real-world web applications. To addr…

Cited by 0SourcecodeScholar
2026

Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders

ICLR 2026poster

Recent multimodal large language models (MLLMs) increasingly integrate multiple vision encoders to improve performance on various benchmarks, assuming that diverse pretraining objectives yield complementary visual signals. However, we show this assumption often fails in practice. Through systematic…

Cited by 0SourcecodeScholar
2025

BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks

EMNLP 2025

Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarcity and ambiguities in mixed-text contexts. We construct English and Chinese Braille Mixed Datasets (EBMD/CBMD) with mat

2025

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

EMNLP 2025

Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand. As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities. However, due to

Cited by 0SourcePDFScholar
2025

ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data

AAAI 2025technical

Recently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particularly after training on document instruction datasets. An effective evaluation method for document instruction data is cruc…

2024

DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing

EMNLP 2024main

Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding. However, previously the research on this topic has been largely hindered since most existing datasets are small-s…

2024

Inversive-Reasoning Augmentation for Natural Language Inference

ICASSP 2024accepted

Natural language inference (NLI) aims to infer the relationship between two texts: premise and hypothesis. However, many existing methods overlook the problem of overestimation of model performance due to superficial correlation biases in NLI datasets. We study this problem and find that most curren…

Cited by 0SourceScholar
2024

LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding

CVPR 2024poster

Recently leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However previous works that employ LLMs/MLLMs for document understanding have not fully explored and utilized the document layout information which…

2023

GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree

EMNLP 2023long main

Inexhaustible web content carries abundant perceptible information beyond text. Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyber-richness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style…

Cited by 0SourceScholar
2023

LORE: Logical Location Regression Network for Table Structure Recognition

AAAI 2023technical

Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they e…

2020

An End-to-End OCR Text Re-organization Sequence Learning for Rich-text Detail Image Comprehension

ECCV 2020poster

Nowadays rich description on detail images help users know more about the commodities. With the help of OCR technology, the description text can be detected and recognized as auxiliary information to remove the comprehending barriers among the visual impaired users. However, for lack of proper logic…

Cited by 29SourcePDFScholar