← Search

Abinay Reddy Naini

8 accepted papers

2026

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools — From Consensus Learning to Ambiguity-Driven Emotion Reasoning

ICML 2026spotlight

Speech Large Language Models (SLLMs) enable high-level emotion reasoning, but often produce ungrounded, text-biased judgments without verifiable acoustic evidence. In contrast, SSL encoders such as WavLM yield strong acoustic representations yet remain opaque discriminative models that offer limited…

Cited by 0SourceScholar
2026

RankList – a Listwise Preference Learning Framework for Predicting Subjective Preferences

AAAI 2026technical

Preference learning has gained significant attention in tasks involving subjective human judgments, such as speech emotion recognition (SER) and image aesthetic assessment. While pairwise frameworks such as RankNet offer robust modeling of relative preferences, they are inherently limited to local c

Cited by 0SourcePDFScholar
2025

Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment

ICASSP 2025accepted

This work addresses the challenge of building speech emotion recognition models that generalize effectively across different domains, particularly when only limited target domain data is available with or without emotional label information. Traditional models often struggle with cross-domain perfor…

Cited by 0SourceScholar
2025

Self-Supervised Learning-Based Multimodal Prediction on Prosocial Behavior Intentions

ICASSP 2025accepted

Human state detection and behavior prediction have seen significant advancements with the rise of machine learning and multimodal sensing technologies. However, predicting prosocial behavior intentions in mobility scenarios, such as helping others on the road, is an underexplored area. Current resea…

Cited by 0SourceScholar
2024

Generalization of Self-Supervised Learning-Based Representations for Cross-Domain Speech Emotion Recognition

ICASSP 2024accepted

Self-supervised learning (SSL) from unlabelled speech data has revolutionized speech representation learning. Among them, wavLM, wav2vec2, HuBERT, and Data2vec have produced benchmark performances on automatic speech recognition. However, few studies have explored the generalization of SSL-based rep…

Cited by 0SourceScholar
2023

Unsupervised Domain Adaptation for Preference Learning Based Speech Emotion Recognition

ICASSP 2023accepted

Retrieving speech samples that have specific expressive content has many applications. It is desirable to build a preference learning framework that ranks speech samples according to emotional attribute values that generalize well to new domains. A popular architecture for preference learning is the…

Cited by 0SourceScholar
2022

Dual Attention Pooling Network for Recording Device Classification Using Neutral and Whispered Speech

ICASSP 2022accepted

In this work, we proposed a method for recording device classification using the recorded speech signal. With the rapid increase in different mobile and professional recording devices, determining the source device has many applications in forensics and in further improving various speech-based appl…

Cited by 0SourceScholar
2019

Formant-gaps Features for Speaker Verification Using Whispered Speech

ICASSP 2019accepted

In this work, we propose a new feature based on formants for whispered speaker verification (SV) task, where neutral data is used for enrollment and whispered recordings are used for test. Such a mismatch between enrollment and test often degrades the performance of whispered SV systems due to the d…

Cited by 0SourceScholar