2025
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
ICLR 2025poster
Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understandi…