← Search

Weichao Zhao

6 accepted papers

2026

DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding

AAAI 2026technical

Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reasoning across pages. While prior work has explored reinforcement learning (RL) for enhancing advanced reasoning in MLLMs, i

Cited by 0SourcePDFScholar
2025

Uni-Sign: Toward Unified Sign Language Understanding at Scale

ICLR 2025poster

Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose…

2024

TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

NeurIPS 2024poster

Tables contain factual and quantitative data accompanied by various structures and contents that pose challenges for machine comprehension. Previous methods generally design task-specific architectures and objectives for individual tasks, resulting in modal isolation and intricate workflows. In this…

2023

BEST: BERT Pre-training for Sign Language Recognition with Coupling Tokenization

AAAI 2023technical

In this work, we are dedicated to leveraging the BERT pre-training success and modeling the domain-specific statistics to fertilize the sign language recognition~(SLR) model. Considering the dominance of hand and body in sign language expression, we organize them as pose triplet units and feed them…

Cited by 37SourcePDFScholar
2021

SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition

ICCV 2021poster

Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In this paper, we introduce the first self-supervised pre-trainable SignBERT with in…

Cited by 104PDFScholar