← Search

Zuan Gao

4 accepted papers

2025

CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

NeurIPS 2025poster

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent benchmarks attempt to address this by focusing on keyword ex…

Cited by 0SourceScholar
2025

SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis

CVPR 2025poster

Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly.To address…

2024

How Control Information Influences Multilingual Text Image Generation and Editing?

NeurIPS 2024poster

Visual text generation has significantly advanced through diffusion models aimed at producing images with readable and realistic text. Recent works primarily use a ControlNet-based framework, employing standard font text images to control diffusion models. Recognizing the critical role of control in…

2024

Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition

IJCAI 2024poster

In text recognition, self-supervised pre-training emerges as a good solution to reduce dependence on expansive annotated real data. Previous studies primarily focus on local visual representation by leveraging mask image modeling or sequence contrastive learning. However, they omit modeling the ling…