← Search

Tanvir Mahmud

4 accepted papers

2024

OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation

EMNLP 2024main

Audio separation in real-world scenarios, where mixtures contain a variable number of sources, presents significant challenges due to limitations of existing models, such as over-separation, under-separation, and dependence on predefined training sources. We propose OpenSep, a novel framework that l…

2024

T-VSL: Text-Guided Visual Sound Source Localization in Mixtures

CVPR 2024poster

Visual sound source localization poses a significant challenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggle to accurately distinguish the semantic regions of each sounding object particu…

2024

Weakly-supervised Audio Separation via Bi-modal Semantic Similarity

ICLR 2024poster

Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop with multi-source training mixtures due to the lack of supervis…

2023

CLIP4VideoCap: Rethinking Clip for Video Captioning with Multiscale Temporal Fusion and Commonsense Knowledge

ICASSP 2023accepted

In this paper, we propose CLIP4VideoCap for video captioning based on large-scale pre-trained CLIP image and text encoders together with multi-scale temporal reasoning and commonsense knowledge. In addition to the CLIP-image encoder operating on successive video frames, we introduce a knowledge dist…

Cited by 0SourceScholar