2026
StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
CVPR 2026
Knowledge-based Visual Question Answering (KVQA) requires models to ground entities in images and reason over factual knowledge. Recent work has introduced its implicit-knowledge variant, IK-KVQA, where a multimodal large language model (MLLM) is the sole knowledge source and answers are produced wi