← Search

Sarthak Yadav

3 accepted papers

2024

Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners

ICLR 2024poster

In this work, we propose a Multi-Window Masked Autoencoder (MW-MAE) fitted with a novel Multi-Window Multi-Head Attention (MW-MHA) module that facilitates the modelling of local-global interactions in every decoder transformer block through attention heads of several distinct local and global window…

Cited by 4SourcePDFScholar
2023

Towards Learning Emotion Information from Short Segments of Speech

ICASSP 2023accepted

Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this appr…

Cited by 0SourceScholar