← Search

Rehana Mahfuz

2 accepted papers

2024

Parameter Efficient Audio Captioning with Faithful Guidance Using Audio-Text Shared Latent Representation

ICASSP 2024accepted

There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models frequently suffer from hallucination and large memory footprint making them challenging to deploy on edge devices. In this pape…

Cited by 0SourceScholar