← Search

Armin Mustafa

11 accepted papers

2025

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

ICLR 2025poster

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models’ ability to generate text descriptions that capture the causal…

2025

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

ICLR 2025poster

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle…

2024

CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing

ECCV 2024poster

"Weakly supervised audio-visual video parsing (AVVP) methods aim to detect audible-only, visible-only, and audible-visible events using only video-level labels. Existing approaches tackle this by leveraging unimodal and cross-modal contexts. However, we argue that while cross-modal learning is benef…

2024

DTF-AT: Decoupled Time-Frequency Audio Transformer for Event Classification

AAAI 2024technical

Convolutional neural networks (CNNs) and Transformer-based networks have recently enjoyed significant attention for various audio classification and tagging tasks following their wide adoption in the computer vision domain. Despite the difference in information distribution between audio spectrogram…

2024

Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event Classification

ICASSP 2024accepted

In the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structure…

Cited by 0SourceScholar
2020

A*3D Dataset: Towards Autonomous Driving in Challenging Environments

ICRA 2020poster

With the increasing global popularity of self-driving cars, there is an immediate need for challenging real-world datasets for benchmarking and training various computer vision tasks such as 3D object detection. Existing datasets either represent simple scenarios or provide only day-time data. In th…

Cited by 206SourcecodeScholar
2016

Temporally Coherent 4D Reconstruction of Complex Dynamic Scenes

CVPR 2016oral

This paper presents an approach for reconstruction of 4D temporally coherent models of complex dynamic scenes. No prior knowledge is required of scene structure or camera calibration allowing reconstruction from multiple moving cameras. Sparse-to-dense temporal correspondence is integrated with join…

Cited by 91PDFScholar