← Search

Bangbang Zhou

3 accepted papers

2025

SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis

CVPR 2025poster

Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly.To address…

2024

Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing

NeurIPS 2024poster

Existing scene text recognition (STR) methods struggle to recognize challenging texts, especially for artistic and severely distorted characters. The limitation lies in the insufficient exploration of character morphologies, including the monotonousness of widely used synthetic training data and the…

2024

Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition

IJCAI 2024poster

Recently, scene text recognition (STR) models have shown significant performance improvements. However, existing models still encounter difficulties in recognizing challenging texts that involve factors such as severely distorted and perspective characters. These challenging texts mainly cause two…