MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
Source attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existing work has primarily focused on text-only scenario and largely overlooked the role of multimodality. We introduce MAVIS,