← Search

Shrikanth S. Narayanan

41 accepted papers

2023

FedAudio: A Federated Learning Benchmark for Audio Tasks

ICASSP 2023accepted

Federated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio…

Cited by 0SourceScholar
2023

Toward Privacy-Enhancing Ambulatory-Based Well-Being Monitoring: Investigating User Re-Identification Risk in Multimodal Data

ICASSP 2023accepted

The sensitivity of data collected via ambulatory monitoring, which regularly involve the recording of speech signals and sensor information, can cause strong privacy concerns. We investigate user re-identification risk in a corpus of such data collected to observe the interplay between behavior, phy…

Cited by 0SourceScholar
2022

Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition

ICASSP 2022accepted

Speech Emotion Recognition (SER) techniques have gained considerable interest in many applications including smart virtual assistants and health state tracking. SER systems often acquire and transmit speech data collected at the client-side to remote cloud platforms for inference and decision making…

Cited by 0SourceScholar
2020

Modeling Behavior as Mutual Dependency between Physiological Signals and Indoor Location in Large-Scale Wearable Sensor Study

ICASSP 2020accepted

Wearable sensors today can unobtrusively collect rich time-series of physiological states and human movement patterns over a prolonged period. Gaining a better understanding of how an individual's physiological responses vary in different workplace environments can be valuable in understanding human…

Cited by 0SourceScholar
2020

Modeling Behavioral Consistency in Large-Scale Wearable Recordings of Human Bio-Behavioral Signals

ICASSP 2020accepted

Continuously-worn wearable sensors provide an unprecedented opportunity to unobtrusively measure rich bio-behavioral time-series recordings in natural settings such as the workplace. These time-series data can be helpful in inferring broad patterns of behavior such as common routines and daily stres…

Cited by 0SourceScholar
2020

Trapezoidal Segment Sequencing: A Novel Approach for Fusion of Human-Produced Continuous Annotations

ICASSP 2020accepted

Generating accurate ground truth representations of human subjective experiences and judgements is essential for advancing our understanding of human-centered constructs such as emotions. Often, this requires the collection and fusion of annotations from several people where each one is subject to v…

Cited by 0SourceScholar
2019

An Empirical Study of Speech Processing in the Brain by Analyzing the Temporal Syllable Structure in Speech-input Induced EEG

ICASSP 2019accepted

Clinical applicability of electroencephalography (EEG) is well established, however the use of EEG as a choice for constructing brain computer interfaces to develop communication platforms is relatively recent. To provide more natural means of communication, there is an increasing focus on bringing…

Cited by 0SourceScholar
2019

Bluetooth Based Indoor Localization Using Triplet Embeddings

ICASSP 2019accepted

We propose a novel algorithm for indoor localization using triplet embeddings through Bluetooth connectivity streams obtained in very noisy settings with irregular sampling schemes using environmental sensors distributed ad hoc inside buildings. We pose the problem as a matrix completion problem, wh…

Cited by 0SourceScholar
2019

Discovering Optimal Variable-length Time Series Motifs in Large-scale Wearable Recordings of Human Bio-behavioral Signals

ICASSP 2019accepted

Continuously-worn wearable sensors produce copious amounts of rich bio-behavioral time series recordings. Exploring recurring patterns, often known as motifs, in wearable time series offers critical insights into understanding the nature of human behavior. Challenges in discovering motifs from weara…

Cited by 0SourceScholar
2019

Improving the Prediction of Therapist Behaviors in Addiction Counseling by Exploiting Class Confusions

ICASSP 2019accepted

In this work we address the problem of joint prosodic and lexical behavioral annotation for addiction counseling. We expand on past work that employed Recurrent Neural Networks (RNNs) on multimodal features by grouping and classifying subsets of classes. We propose two implementations: One is hierar…

Cited by 0SourceScholar
2019

Learning Shared Vector Representations of Lyrics and Chords in Music

ICASSP 2019accepted

Music has a powerful influence on a listener's emotions. In this paper, we represent lyrics and chords in a shared vector space using a phrase-aligned chord-and-lyrics corpus. We show that models that use these shared representations predict a listener's emotion while hearing musical passages better…

Cited by 0SourceScholar
2019

On Evaluating CNN Representations for Low Resource Medical Image Classification

ICASSP 2019accepted

Convolutional Neural Networks (CNNs) have revolutionized performances in several machine learning tasks such as image classification, object tracking, and keyword spotting. However, given that they contain a large number of parameters, their direct applicability into low resource tasks is not straig…

Cited by 0SourceScholar
2019

On Role and Location of Normalization before Model-based Data Augmentation in Residual Blocks for Classification Tasks

ICASSP 2019accepted

Regularization is crucial to the success of many practical deep learning models, in particular in frequent scenarios where there are only a few to a moderate number of accessible training samples. In addition to weight decay, noise injection and dropout, regularization based on multi-branch architec…

Cited by 0SourceScholar
2019

Reinforcing Self-expressive Representation with Constraint Propagation for Face Clustering in Movies

ICASSP 2019accepted

The ability to robustly cluster faces in movies is a necessary step in understanding media content representations of people along dimensions such as gender and age. Building upon the successes of sparse subspace clustering (SSC) in uncovering the underlying structure of the data, in this paper we p…

Cited by 0SourceScholar
2019

Robust Speech Activity Detection in Movie Audio: Data Resources and Experimental Evaluation

ICASSP 2019accepted

Speech activity detection in highly variable acoustic conditions is a challenging task. Many approaches to detect speech activity in such conditions involve an inherent knowledge of the noise types involved. Movie audio can offer an excellent research test-bed for developing speech activity models.…

Cited by 0SourceScholar
2019

Role Specific Lattice Rescoring for Speaker Role Recognition from Speech Recognition Outputs

ICASSP 2019accepted

The language patterns followed by different speakers who play specific roles in conversational interactions provide valuable cues for the task of Speaker Role Recognition (SRR). Given the speech signal, existing algorithms typically try to find such patterns in the output of an Automatic Speech Reco…

Cited by 0SourceScholar
2019

Speaker Agnostic Foreground Speech Detection from Audio Recordings in Workplace Settings from Wearable Recorders

ICASSP 2019accepted

Audio-signal acquisition as part of wearable sensing adds an important dimension for applications such as understanding human behaviors. As part of a large study on work place behaviours, we collected audio data from individual hospital staff using custom wearable recorders. The audio features colle…

Cited by 0SourceScholar
2019

Toward Robust Interpretable Human Movement Pattern Analysis in a Workplace Setting

ICASSP 2019accepted

Gaining a better understanding of how people move about and interact with their environment is an important piece of understanding human behavior. Careful analysis of individuals' deviations or variations in movement over time can provide an awareness about changes to their physical or mental state…

Cited by 0SourceScholar
2018

A Novel Method for Human Bias Correction of Continuous- Time Annotations

ICASSP 2018accepted

Human annotations are of integral value in human behavior studies and in particular for the generation of ground truth for behavior prediction using various machine learning methods. These often subjective human annotations are especially required for studies involving measuring and predicting hidde…

Cited by 0SourceScholar
2018

Improving Semi-Supervised Classification for Low-Resource Speech Interaction Applications

ICASSP 2018accepted

We propose a semi-supervised learning method to improve classification performance in scenarios with limited labeled data. We employ adaptation strategies such as entropy-filtering and self-training, and show that our method achieves up to 17.2% relative improvement in UAR for a multi-class problem.…

Cited by 0SourceScholar
2018

Semi-Supervised and Transfer Learning Approaches for Low Resource Sentiment Classification

ICASSP 2018accepted

Sentiment classification involves quantifying the affective reaction of a human to a document, media item or an event. Although researchers have investigated several methods to reliably infer sentiment from lexical, speech and body language cues, training a model with a small set of labeled datasets…

Cited by 0SourceScholar
2018

Shaking Acoustic Spectral Sub-Bands can Letxer Regularize Learning in Affective Computing

ICASSP 2018accepted

In this work, we investigate a recently proposed regularization technique based on multi-branch architectures, called Shake-Shake regularization, for the task of speech emotion recognition. In addition, we also propose variants to incorporate domain knowledge into model configurations. The experimen…

Cited by 0SourceScholar
2017

A knowledge transfer and boosting approach to the prediction of affect in movies

ICASSP 2017accepted

Affect prediction is a classical problem and has recently garnered special interest in multimedia applications. Affect prediction in movies is one such domain, potentially aiding the design as well as the impact analysis of movies. Given the large diversity in movies (such as different genres and la…

Cited by 0SourceScholar
2017

A knowledge-driven framework for ECG representation and interpretation for wearable applications

ICASSP 2017accepted

The increasing use of wearable technology creates the need for reliable signal representations with low storage and transmission cost, as well as interpretable models that can be used to translate signals into meaningful constructs. We propose a knowledge-driven sparse representation of the electroc…

Cited by 0SourceScholar
2017

Estimation of vocal tract area function from volumetric Magnetic Resonance Imaging

ICASSP 2017accepted

The acoustic properties of speech signals are largely determined by the shaping of the vocal tract. Thus, measurements of vocal-tract area functions and their relationship to various properties of the speech signal have been of interest to the speech research community. Recent advances in Magnetic R…

Cited by 0SourceScholar
2017

Quantifying regulation mechanisms in dating couples through a dynamical systems model of acoustic and physiological arousal

ICASSP 2017accepted

Negative emotional arousal during conflict has been related to negative outcomes in romantic relationships and degraded quality of family life. Despite its extensive study in psychology, it is still challenging to quantify emotional arousal in a meaningful way with objective indices beyond tradition…

Cited by 0SourceScholar
2017

Towards a definition of local stationarity for graph signals

ICASSP 2017accepted

In this paper, we extend the recent definition of graph stationarity into a definition of local stationarity. Doing so, we present a metric to assess local stationarity using projections on localized atoms on the graph. Energy of these projections defines the local power spectrum of the signal. We u…

Cited by 0SourceScholar
2016

A multimodal mixture-of-experts model for dynamic emotion prediction in movies

ICASSP 2016accepted

This paper addresses the problem of continuous emotion prediction in movies from multimodal cues. The rich emotion content in movies is inherently multimodal, where emotion is evoked through both audio (music, speech) and video modalities. To capture such affective information, we put forth a set of…

Cited by 0SourceScholar
2016

CNMF-based acoustic features for noise-robust ASR

ICASSP 2016accepted

We present an algorithm using convolutive non-negative matrix factorization (CNMF) to create noise-robust features for automatic speech recognition (ASR). Typically in noise-robust ASR, CNMF is used to remove noise from noisy speech prior to feature extraction. However, we find that denoising introd…

Cited by 0SourceScholar
2016

Lightly-supervised utterance-level emotion identification using latent topic modeling of multimodal words

ICASSP 2016accepted

Research on multimodal emotion recognition has drawn much attention recently in diverse disciplines. With the increasing amount of multimodal data, unsupervised or semi-supervised learning has become highly desirable to automatically discover expression of emotion patterns in behavioral data. We pre…

Cited by 0SourceScholar
2016

Opening big in box office? Trailer content can help

ICASSP 2016accepted

Computational prediction of a movie's financial success usually relies only on metadata such as - genre, budget, actors, Motion Picture Association of America (MPAA) rating and critics' reviews. We argue that movie trailers, created to invoke viewers' interest and curiosity about a movie, carry comp…

Cited by 0SourceScholar
2016

Pathological speech processing: State-of-the-art, current challenges, and future directions

ICASSP 2016accepted

The study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP…

Cited by 0SourceScholar
2015

A mixture of experts approach towards intelligibility classification of pathological speech

ICASSP 2015accepted

Pathological speech involves atypical speech production which may result from several factors including oral diseases, physical disabilities in the voice production system and atypical anatomy. Automatic evaluation of intelligibility in patients with pathological speech can assist accurate diagnosis…

Cited by 0SourceScholar
2015

Computationally deconstructing movie narratives: An informatics approach

ICASSP 2015accepted

In general, popular films and screenplays follow a well defined storytelling paradigm that comprises three essential segments or acts: exposition (act I), conflict (act II) and resolution (act III). Deconstructing a movie into its narrative units can enrich semantic understanding of movies, and help…

Cited by 0SourceScholar
2015

Improvements to the IBM speech activity detection system for the DARPA RATS program

ICASSP 2015accepted

In this paper we describe improvements to the IBM speech activity detection (SAD) system for the third phase of the DARPA RATS program. The progress during this final phase comes from jointly training convolutional and regular deep neural networks with rich time-frequency representations of speech.…

Cited by 0SourceScholar
2015

Modeling mutual influence of multimodal behavior in affective dyadic interactions

ICASSP 2015accepted

To accomplish effective communication, interaction partners generally adapt their verbal and non-verbal behavior to that of their interlocutors. This behavior adaptation is often modulated by the underlying emotional states of partners. Modeling such mutual behavioral influence is critical for emoti…

Cited by 0SourceScholar
2015

On quantifying facial expression-related atypicality of children with Autism Spectrum Disorder

ICASSP 2015accepted

by adult observers. This paper focuses on data driven ways to analyze and quantify atypicality in facial expressions of children with ASD. Our objective is to uncover those characteristics of facial gestures that induce the sense of perceived atypicality in observers. Using a carefully collected mot…

Cited by 0SourceScholar
2015

Quantifying EDA synchrony through joint sparse representation: A case-study of couples' interactions

ICASSP 2015accepted

The co-variation degree between individuals in their physiological signals can reveal insights about the quality of their interaction as well as their personal characteristics. In an effort to capture the amount of synchrony between Electrodermal Activity (EDA) streams occurring in parallel during d…

Cited by 0SourceScholar
2015

Redundancy analysis of behavioral coding for couples therapy and improved estimation of behavior from noisy annotations

ICASSP 2015accepted

Assessment and quantification of behavior is an important research objective in the recently developed field of behavioral signal processing. This paper focuses on the estimation of behavior from noisy human assessment. It aims to address the redundancy of behavioral descriptors for couples therapy…

Cited by 0SourceScholar