← Search

Shiva Sundaram

9 accepted papers

2023

Multi-Scale Compositional Constraints for Representation Learning on Videos

ICASSP 2023accepted

Combining simple concepts to form structured thoughts and decomposing complex concepts into their constituents is one key characteristic of human cognition. In this work we extract video representations by combining multi-scale processing with compositional constraints, i.e., we constrain the latent…

Cited by 0SourceScholar
2022

Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation

ICASSP 2022accepted

Audio-visual data allows us to leverage different modalities for downstream tasks. The idea being individual streams can complement each other in the given task, thereby resulting in a model with improved performance. In this work, we present our experimental results on action recognition and video…

Cited by 0SourceScholar
2021

Disentanglement for Audio-Visual Emotion Recognition Using Multitask Setup

ICASSP 2021accepted

Deep learning models trained on audio-visual data have been successfully used to achieve state-of-the-art performance for emotion recognition. In particular, models trained with multitask learning have shown additional performance improvements. However, such multitask models entangle information bet…

Cited by 0SourceScholar
2020

Fully Learnable Front-End for Multi-Channel Acoustic Modeling Using Semi-Supervised Learning

ICASSP 2020accepted

In this work, we investigated the teacher-student training paradigm to train a fully learnable multi-channel acoustic model for far-field automatic speech recognition (ASR). Using a large offline teacher model trained on beamformed audio, we trained a simpler multi-channel student acoustic model use…

Cited by 0SourceScholar
2020

Robust Multi-Channel Speech Recognition Using Frequency Aligned Network

ICASSP 2020accepted

Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial filtering layer jointly within an acoustic model. In this paper,…

Cited by 0SourceScholar
2019

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

ICASSP 2019accepted

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques do not always yield ASR accuracy improvement because the op…

Cited by 39SourceScholar
2019

Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning

ICASSP 2019accepted

For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we a…

Cited by 53SourceScholar
2019

Multi-geometry Spatial Acoustic Modeling for Distant Speech Recognition

ICASSP 2019accepted

The use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array geometry mismatch between design and test conditions. Moreover,…

Cited by 18SourceScholar