← Search

Amit Namburi

3 accepted papers

2025

CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation

ICASSP 2025accepted

Modeling temporal characteristics plays a significant role in the representation learning of audio waveform. We propose Contrastive Long-form Language-Audio Pretraining (CoLLAP) to significantly extend the perception window for both the input audio (up to 5 minutes) and the language descriptions (ex…

Cited by 0SourceScholar
2025

FUTGA-MIR: Enhancing Fine-grained and Temporally-aware Music Understanding with Music Information Retrieval

ICASSP 2025accepted

Recent music large language models (music LLMs) have shown great potential in music understanding through large-scale multimodal pre-training. While some existing music LLMs have been augmented with temporally-aware music captions, music information retrieval (MIR) features conventionally do not exi…

Cited by 0SourceScholar
2025

WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning

EMNLP 2025

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remain largely unexplored.We introduce WildScore, the first in-the-wild multimodal sy