← Search

Tony Alex

4 accepted papers

2026

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

ICML 2026poster

Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA…

Cited by 0SourceScholar
2025

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

ICLR 2025poster

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle…

2024

DTF-AT: Decoupled Time-Frequency Audio Transformer for Event Classification

AAAI 2024technical

Convolutional neural networks (CNNs) and Transformer-based networks have recently enjoyed significant attention for various audio classification and tagging tasks following their wide adoption in the computer vision domain. Despite the difference in information distribution between audio spectrogram…

2024

Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event Classification

ICASSP 2024accepted

In the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structure…

Cited by 0SourceScholar