Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
Large Vision-Language Models (LVLMs) have experienced significant advancements in recent years. However, their performance still falls short in tasks requiring deep visual perception, such as identifying subtle differences between images. A potential cause is the scarcity of visual knowledge in popu