LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from eac…