Spice+: Evaluation of Automatic Audio Captioning Systems with Pre-Trained Language Models
Audio captioning aims at describing acoustic scenes with natural language. Systems are currently evaluated by image captioning metrics CIDEr and SPICE. However, recent studies have highlighted a poor correlation of these metrics with human assessments. In this paper, we propose SPICE+, a modificatio…