← Search

Sara Ahmed

3 accepted papers

2024

DTF-AT: Decoupled Time-Frequency Audio Transformer for Event Classification

AAAI 2024technical

Convolutional neural networks (CNNs) and Transformer-based networks have recently enjoyed significant attention for various audio classification and tagging tasks following their wide adoption in the computer vision domain. Despite the difference in information distribution between audio spectrogram…

2024

Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event Classification

ICASSP 2024accepted

In the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structure…

Cited by 0SourceScholar
2024

SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action Recognition

AAAI 2024technical

Contrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel cont…