2025
SiQA: A Large Multi-Modal Question Answering Model for Structured Images Based on RAG
ICASSP 2025accepted
Existing Large Multimodal Models (LMMs) demonstrate excellent performance in handling visual tasks in everyday scenarios. However, they still face challenges in understanding structured images, such as flowcharts and organizational charts, which are characterized by text-rich and complex hierarchica…