SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis
Due to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly.To address…