← Search

Shrikanth Narayanan

44 accepted papers

2026

ENCODING EMOTION THROUGH SELF-SUPERVISED EYE MOVEMENT RECONSTRUCTION

ICASSP 2026poster

The relationship between emotional expression and eye movement is well-documented, with literature establishing gaze patterns are reliable indicators of emotion. However, most studies utilize specialized, high-resolution eye-tracking equipment, limiting the potential reach of findings. We investigat…

Cited by 0SourcePDFScholar
2025

Aggregation Artifacts in Subjective Tasks Collapse Large Language Models’ Posteriors

NAACL 2025long

In-context Learning (ICL) has become the primary method for performing natural language tasks with Large Language Models (LLMs). The knowledge acquired during pre-training is crucial for this few-shot capability, providing the model with task priors. However, recent studies have shown that ICL predo…

2025

Creating a Lens of Chinese Culture: A Multimodal Dataset for Chinese Pun Rebus Art Understanding

ACL 2025finding

Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly culturally rich art forms, remains less explored. As a pearl of human wisdom and creativity, art encapsulates complex cultural n…

2025

Data Efficient Child-Adult Speaker Diarization with Simulated Conversations

ICASSP 2025accepted

Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies "who spoke when", is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privac…

Cited by 0SourceScholar
2025

Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction

ICASSP 2025accepted

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We…

Cited by 0SourceScholar
2025

Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts

EMNLP 2025

Modeling complex subjective tasks in Natural Language Processing, such as recognizing emotion and morality, is considerably challenging due to significant variation in human annotations. This variation often reflects reasonable differences in semantic interpretations rather than mere noise, necessit

2025

Large Language Models Do Multi-Label Classification Differently

EMNLP 2025

Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied. We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the mode

2025

Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks

ICASSP 2025accepted

In-Context Learning (ICL) in Large Language Models (LLM) has emerged as the dominant technique for performing natural language tasks, as it does not require updating the model parameters with gradient-based methods. ICL promises to "adapt" the LLM to perform the present task at a competitive or stat…

Cited by 0SourceScholar
2025

Scaling Wearable Foundation Models

ICLR 2025poster

Wearable sensors have become ubiquitous thanks to a variety of health tracking features. The resulting continuous and longitudinal measurements from everyday life generate large volumes of data. However, making sense of these observations for scientific and actionable insights is non-trivial. Inspir…

Cited by 6SourcePDFScholar
2025

Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech

ICASSP 2025accepted

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Reson…

Cited by 0SourceScholar
2025

Wavelet Scattering Network Features for Intensity Category Classification and Prediction of SPL from Speech

ICASSP 2025accepted

Speakers change vocal intensity in daily life to communicate over long distances and to express vocal emotions. Humans produce speech using different intensity categories (e.g. soft, normal and loud voice) and they can regulate intensity across a wide sound pressure level (SPL) range. Knowing the in…

Cited by 0SourceScholar
2024

Audio-Visual Child-Adult Speaker Classification in Dyadic Interactions

ICASSP 2024accepted

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer scale and robustness across diverse and wide-ranging conditions.…

Cited by 0SourceScholar
2024

CVAT-BWV: A Web-Based Video Annotation Platform for Police Body-Worn Video

IJCAI 2024poster

We introduce an open-source platform for annotating body-worn video (BWV) footage aimed at enhancing transparency and accountability in policing. Despite the widespread adoption of BWVs in police departments, analyzing the vast amount of footage generated has presented significant challenges. This i…

2024

Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization

ICASSP 2024accepted

Video summarization remains a huge challenge in computer vision due to the size of the input videos to be summarized. We propose an efficient, language-only video summarizer that achieves competitive accuracy with high data efficiency. Using only textual captions obtained via a zero-shot approach, w…

Cited by 0SourceScholar
2024

Emotion-Aligned Contrastive Learning Between Images and Music

ICASSP 2024accepted

Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using queries of various modalities including text, video, and spe…

Cited by 0SourceScholar
2024

Foundation Model Assisted Automatic Speech Emotion Recognition: Transcribing, Annotating, and Augmenting

ICASSP 2024accepted

Significant advances are being made in speech emotion recognition (SER) using deep learning models. Nonetheless, training SER systems remains challenging, requiring both time and costly resources. Like many other machine learning tasks, acquiring datasets for SER requires substantial data annotation…

Cited by 0SourceScholar
2024

TRUST-SER: On The Trustworthiness Of Fine-Tuning Pre-Trained Speech Embeddings For Speech Emotion Recognition

ICASSP 2024accepted

Recent studies have explored using pre-trained embeddings for speech emotion recognition, achieving comparable performance to conventional methods that rely on low-level knowledge-inspired acoustic features. These embeddings are often generated from models trained on large-scale speech datasets usin…

Cited by 0SourceScholar
2023

A Context-Aware Computational Approach for Measuring Vocal Entrainment in Dyadic Conversations

ICASSP 2023accepted

Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual’s cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal entrainment in dyadic conversations. We use conformers (a combination…

Cited by 0SourceScholar
2023

A Dataset for Audio-Visual Sound Event Detection in Movies

ICASSP 2023accepted

Audio event detection is a widely studied field, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically involve manual annotation and verification, which is expensive to perform at s…

Cited by 0SourceScholar
2023

Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?

ICASSP 2023accepted

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-bas…

Cited by 0SourceScholar
2023

Contextually-Rich Human Affect Perception Using Multimodal Scene Information

ICASSP 2023accepted

The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived…

Cited by 0SourceScholar
2023

Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAP

ICASSP 2023accepted

There is an imminent need for guidelines and standard test sets to allow direct and fair comparisons of speech emotion recognition (SER). While resources, such as the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, have emerged as widely-adopted reference corpora for researchers to d…

Cited by 0SourceScholar
2023

Domain Adaptation for Sentiment Analysis Using Robust Internal Representations

EMNLP 2023long findings

Sentiment analysis is a costly yet necessary task for enterprises to study the opinions of their customers to improve their products and to determine optimal marketing strategies. Due to the existence of a wide range of domains across different products and services, cross-domain sentiment anal…

Cited by 0SourceScholar
2023

Leveraging Label Correlations in a Multi-Label Setting: a Case Study in Emotion

ICASSP 2023accepted

Detecting emotions expressed in text has become critical to a range of fields. In this work, we investigate ways to exploit label correlations in multi-label emotion recognition models to improve emotion detection. First, we develop two modeling approaches to the problem in order to capture word ass…

Cited by 0SourceScholar
2023

Navigating and Reaching Therapeutic Goals with Dynamical Systems in Conversation-Based Interventions

ICASSP 2023accepted

Modern human behavioral signal processing and machine-learning methods have introduced novel ways for representing and estimating internal states of people in goal-based conversational interactions, such as psychotherapy. By combining these methods with systems theoretic approaches, we demonstrate h…

Cited by 0SourceScholar
2023

On the Role of Visual Context in Enriching Music Representations

ICASSP 2023accepted

Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval. Incorporating multimodal context from diverse sources provides a pro…

Cited by 0SourceScholar
2023

Signal Processing Grand Challenge 2023 - E-Prevention: Sleep Behavior as an Indicator of Relapses in Psychotic Patients

ICASSP 2023accepted

This paper presents the approach and results of USC SAIL’s submission to the Signal Processing Grand Challenge 2023 – e-Prevention (Task 2), on detecting relapses in psychotic patients. Relapse prediction has proven to be challenging, primarily due to the heterogeneity of symptoms and responses to t…

Cited by 0SourceScholar
2023

Using Emotion Embeddings to Transfer Knowledge between Emotions, Languages, and Annotation Formats

ICASSP 2023accepted

The need for emotional inference from text continues to diversify as more and more disciplines integrate emotions into their theories and applications. These needs include inferring different emotion types, handling multiple languages, and different annotation formats. A shared model between differe…

Cited by 0SourceScholar
2022

Leveraging Open Data and Task Augmentation to Automated Behavioral Coding of Psychotherapy Conversations in Low-Resource Scenarios

EMNLP 2022finding

In psychotherapy interactions, the quality of a session is assessed by codifying the communicative behaviors of participants during the conversation through manual observation and annotation. Developing computational approaches for automated behavioral coding can reduce the burden on human coders an…

2021

Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training

ICASSP 2021accepted

Deep neural network based speaker recognition systems can easily be deceived by an adversary using minuscule imperceptible perturbations to the input speech samples. These adversarial attacks pose serious security threats to the speaker recognition systems that use speech biometric. To address this…

Cited by 0SourceScholar
2021

Context-Aware Speech Stress Detection in Hospital Workers Using Bi-LSTM Classifiers

ICASSP 2021accepted

Hospital workers are known to work long hours in a highly stressful environment. The COVID-19 pandemic has increased this burden multi-fold. Pre-COVID statistics already showed that one in every three nurses reported burnout, thus affecting patient satisfaction and the quality of their provided serv…

Cited by 0SourceScholar
2020

Automatic Prediction of Suicidal Risk in Military Couples Using Multimodal Interaction Cues from Couples Conversations

ICASSP 2020accepted

Suicide is a major societal challenge globally, with a wide range of risk factors, from individual health, psychological and behavioral elements to socio-economic aspects. Military personnel, in particular, are at especially high risk. Crisis resources, while helpful, are often constrained by access…

Cited by 0SourceScholar
2020

Bringing in the Outliers: A Sparse Subspace Clustering Approach to Learn a Dictionary of Mouse Ultrasonic Vocalizations

ICASSP 2020accepted

Mice vocalize in the ultrasonic range during social interactions. These vocalizations are used in neuroscience and clinical studies to tap into complex behaviors and states. The analysis of these ultrasonic vocalizations (USVs) has been traditionally a manual process, which is prone to errors and hu…

Cited by 0SourceScholar
2020

Identifying Truthful Language in Child Interviews

ICASSP 2020accepted

When a child is suspected to be the victim or sole witness of a crime, the manner in which information is gathered from the child becomes critical. A child forensic interview is the guided conversation that a legal expert conducts to elicit reliable information from a child. To help substantiate chi…

Cited by 0SourceScholar
2020

Learning Domain Invariant Representations for Child-Adult Classification from Speech

ICASSP 2020accepted

Diagnostic procedures for ASD (autism spectrum disorder) involve semi-naturalistic interactions between the child and a clinician. Computational methods to analyze these sessions require an end-to-end speech and language processing pipeline that goes from raw audio to clinically-meaningful behaviora…

Cited by 0SourceScholar
2020

Meta-Learning for Robust Child-Adult Classification from Speech

ICASSP 2020accepted

Computational modeling of naturalistic conversations in clinical applications has seen growing interest in the past decade. An important use-case involves child-adult interactions within the autism diagnosis and intervention domain. In this paper, we address a specific sub-problem of speaker diariza…

Cited by 0SourceScholar
2020

Robust Speaker Recognition Using Unsupervised Adversarial Invariance

ICASSP 2020accepted

In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker emb…

Cited by 0SourceScholar
2020

Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network

ICASSP 2020accepted

In this work, we propose deep latent space clustering for speaker diarization using generative adversarial network (GAN) back-projection with the help of an encoder network. The proposed diarization system is trained jointly with GAN loss, latent variable recovery loss, and a clustering-specific los…

Cited by 0SourceScholar
2020

Speaker-Invariant Affective Representation Learning via Adversarial Training

ICASSP 2020accepted

Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold-standard references. In addition, there is much variability from input speech signals, human subjective perception of the signals and emotion label ambiguity. In this paper, we p…

Cited by 0SourceScholar
2020

The Role of Annotation Fusion Methods in the Study of Human-Reported Emotion Experience During Music Listening

ICASSP 2020accepted

Music is a universally-enjoyed art form, but listeners often respond to it in tremendously different ways. The same song can bring one person great joy and another deep sorrow. This paper focuses on modeling human music experience at the group level. In this scenario, human annotations serve an impo…

Cited by 0SourceScholar
2020

Vocal Tract Articulatory Contour Detection in Real-Time Magnetic Resonance Images Using Spatio-Temporal Context

ICASSP 2020accepted

Due to its ability to visualize and measure the dynamics of vocal tract shaping during speech production, real-time magnetic resonance imaging (rtMRI) has emerged as one of the prominent research tools. The ability to track different articulators such as the tongue, lips, velum, and the pharynx is a…

Cited by 0SourceScholar