← Search

Cha Zhang

12 accepted papers

2025

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

ICML 2025poster

Structured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selec…

Cited by 4SourcePDFScholar
2023

From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding

ACL 2023long

Current state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens. This process known as tokenization relies on a pre-built vocabulary of words or sub-word morphemes. This fixed vocabulary limits the model’s robustness to spellin…

Cited by 8SourcePDFScholar
2023

TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models

AAAI 2023technical

Text recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-process…

2023

Unifying Vision, Text, and Layout for Universal Document Processing

CVPR 2023highlight

We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to mo…

2022

XDoc: Unified Pre-training for Cross-Format Document Understanding

EMNLP 2022finding

The surge of pre-training has witnessed the rapid development of document understanding recently. Pre-training and fine-tuning framework has been effectively used to tackle texts in various formats, including plain texts, document texts, and web texts. Despite achieving promising performance, existi…

2022

XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding

ACL 2022findings

Multimodal pre-training with text, layout, and image has achieved SOTA performance for visually rich document understanding tasks recently, which demonstrates the great potential for joint learning across different modalities. However, the existed research work has focused only on the English domain…

2021

LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding

ACL 2021long

Pre-training of text and layout has proved effective in a variety of visually-rich document understanding tasks due to its effective model architecture and the advantage of large-scale unlabeled scanned/digital-born documents. We propose LayoutLMv2 architecture with new pre-training tasks to model t…

2021

TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption

CVPR 2021poster

In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to the conventional vision-language pre-training that fail…

Cited by 192PDFcodeScholar
2020

Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing

ICASSP 2020accepted

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a…

Cited by 0SourceScholar
2020

Towards Efficient Model Compression via Learned Global Ranking

CVPR 2020oral

Pruning convolutional filters has demonstrated its effectiveness in compressing ConvNets. Prior art in filter pruning requires users to specify a target model complexity (e.g., model size or FLOP count) for the resulting architecture. However, determining a target model complexity can be difficult f…

Cited by 232PDFcodeScholar
2017

Automatic speech emotion recognition using recurrent neural networks with local attention

ICASSP 2017accepted

Automatic emotion recognition from speech is a challenging task which relies heavily on the effectiveness of the speech features used for classification. In this work, we study the use of deep learning to automatically discover emotionally relevant features from speech. It is shown that using a deep…

Cited by 0SourceScholar