2024
Parameter Efficient Audio Captioning with Faithful Guidance Using Audio-Text Shared Latent Representation
ICASSP 2024accepted
There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models frequently suffer from hallucination and large memory footprint making them challenging to deploy on edge devices. In this pape…