← Search

Rajiv Ratn Shah

35 accepted papers

2026

ALPHA: Action-Based Learning for Pluralistic Human Alignment in Large Language Models

AAAI 2026technical

Large language models are widely used, but aligning them with societal values remains challenging. Current approaches often rely on human annotations, which are hard to scale, or synthetic data produced by models that may themselves be misaligned, making it difficult to capture genuine public opinio

Cited by 0SourcePDFScholar
2026

BRI-MH: Behavioral Risk Index for Mental Health — An Interpretable Multimodal LLM-Augmented Framework (Student Abstract)

AAAI 2026technical

Mental health monitoring faces challenges from fragmented data and opaque risk scores. We present BRI-MH, an in- terpretable multimodal framework combining behavioral sig- nals with cognitive features from large language models to produce a weekly Behavioral Risk Index. Unlike prior work with isolat

Cited by 0SourcePDFScholar
2026

IMPACT: Integrated Multimodal Pipeline for Rapid Accident Causality Tracking (Student Abstract)

AAAI 2026technical

Traffic accidents pose a significant societal challenge, with many fatalities being avoidable through timely emergency response. We introduce IMPACT (Integrated Multimodal Pipeline for Rapid Accident Causality Tracking), a scalable AI framework designed for autonomous, rapid traffic incident analysi

Cited by 0SourcePDFScholar
2026

Social Agents: Collective Intelligence Improves LLM Predictions

ICLR 2026poster

In human society, collective decision making has often outperformed the judgment of individuals. Classic examples range from estimating livestock weights to predicting elections and financial markets, where averaging many independent guesses often yields results more accurate than experts. These suc…

Cited by 0SourceScholar
2026

When Equal Isn’t Fair: Mitigating Over-Normalization in Large Language Models (Student Abstract)

AAAI 2026technical

Bias in Large Language Models (LLMs) is increasingly addressed through fairness-oriented techniques. However, in some cases, these approaches may inadvertently remove genuine cultural differences between groups, leading to “over-normalization” or models losing important socio-cultural distinctions.

Cited by 0SourcePDFScholar
2025

CompMTL: Layer-Wise Competitive Multi-Task Learning

ICASSP 2025accepted

It is challenging to simultaneously address multiple related tasks using a unified multi-task model and consistently balance conflicts across these tasks. The conflicts arise because each task competes to update the shared module in a manner that can better align with its own requirements. To addres…

Cited by 0SourceScholar
2025

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion

AAAI 2025technical

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion intensity in the diffusion-based EVC framework to generate precise…

Cited by 0SourcePDFScholar
2025

Measuring And Improving Engagement of Text-to-Image Generation Models

ICLR 2025poster

Recent advances in text-to-image generation have achieved impressive aesthetic quality, making these models usable for both personal and commercial purposes. However, in the fields of marketing and advertising, images are often created to be more engaging, as reflected in user behaviors such as incr…

2025

Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English

AAAI 2025technical

Large Language Models (LLMs) excel in linguistic tasks but struggle with mathematical reasoning, particularly in non- English languages like Hindi. This research aims to en- hance the mathematical reasoning skills of smaller, resource- efficient open-source LLMs in both Hindi and English. We evaluat…

2025

SPRO: Improving Image Generation via Self-Play

NeurIPS 2025poster

Recent advances in diffusion models have dramatically improved image fidelity and diversity. However, aligning these models with nuanced human preferences -such as aesthetics, engagement, and subjective appeal remains a key challenge due to the scarcity of large-scale human annotations. Collecting s…

Cited by 0SourceScholar
2025

Teaching Human Behavior Improves Content Understanding Abilities Of VLMs

ICLR 2025poster

Communication is defined as "*Who* says *what* to *whom* with *what* effect." A message from a communicator generates downstream receiver effects, also known as behavior. Receiver behavior, being a downstream effect of the message, carries rich signals about it. Even after carrying signals about the…

2024

Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior

ICLR 2024spotlight

Shannon and Weaver's seminal information theory divides communication into three levels: technical, semantic, and effectiveness. While the technical level deals with the accurate reconstruction of transmitted symbols, the semantic and effectiveness levels deal with the inferred meaning and its effec…

2024

Multilingual Coreference Resolution in Low-resource South Asian Languages

COLING 2024main

Coreference resolution involves the task of identifying text spans within a discourse that pertain to the same real-world entity. While this task has been extensively explored in the English language, there has been a notable scarcity of publicly accessible resources and models for coreference resol…

2023

A Video Is Worth 4096 Tokens: Verbalize Story Videos To Understand Them In Zero Shot

EMNLP 2023long main

Multimedia content, such as advertisements and story videos, exhibit a rich blend of creativity and multiple modalities. They incorporate elements like text, visuals, audio, and storytelling techniques, employing devices like emotions, symbolism, and slogans to convey meaning. There is a dearth of l…

Cited by 0SourceScholar
2023

Analysing the Masked Predictive Coding Training Criterion for Pre-Training a Speech Representation Model

ICASSP 2023accepted

Recent developments in pre-trained speech representation utilizing self-supervised learning (SSL) have yielded exceptional results on a variety of downstream tasks. One such technique, known as masked predictive coding (MPC), has been employed by some of the most high-performing models. In this stud…

Cited by 0SourceScholar
2023

H-AES: Towards Automated Essay Scoring for Hindi

AAAI 2023technical

The use of Natural Language Processing (NLP) for Automated Essay Scoring (AES) has been well explored in the English language, with benchmark models exhibiting performance comparable to human scorers. However, AES in Hindi and other low-resource languages remains unexplored. In this study, we reprod…

2023

Mask-Net: Learning Context Aware Invariant Features Using Adversarial Forgetting (Student Abstract)

AAAI 2023technical

Training a robust system, e.g., Speech to Text (STT), requires large datasets. Variability present in the dataset, such as unwanted nuances and biases, is the reason for the need for large datasets to learn general representations. In this work, we propose a novel approach to induce invariance using…

2023

Persuasion Strategies in Advertisements

AAAI 2023technical

Modeling what makes an advertisement persuasive, i.e., eliciting the desired response from consumer, is critical to the study of propaganda, social psychology, and marketing. Despite its importance, computational modeling of persuasion in computer vision is still in its infancy, primarily due to the…

2022

MINIMAL: Mining Models for Universal Adversarial Triggers

AAAI 2022technical

It is well known that natural language models are vulnerable to adversarial attacks, which are mostly input-specific in nature. Recently, it has been shown that there also exist input-agnostic attacks in NLP models, called universal adversarial triggers. However, existing methods to craft universal…

Cited by 5SourcePDFScholar
2021

An Empirical Investigation of Bias in the Multimodal Analysis of Financial Earnings Calls

NAACL 2021long

Volatility prediction is complex due to the stock market’s stochastic nature. Existing research focuses on the textual elements of financial disclosures like earnings calls transcripts to forecast stock volatility and risk, but ignores the rich acoustic features in the company executives’ speech. Re…

2021

Enhanced Audio Tagging via Multi- to Single-Modal Teacher-Student Mutual Learning

AAAI 2021technical

Recognizing ongoing events based on acoustic clues has been a critical yet challenging problem that has attracted significant research attention in recent years. Joint audio-visual analysis can improve the event detection accuracy but may not always be feasible as under many circumstances only audio…

Cited by 16SourcePDFScholar
2021

GupShup: Summarizing Open-Domain Code-Switched Conversations

EMNLP 2021main

Code-switching is the communication phenomenon where the speakers switch between different languages during a conversation. With the widespread adoption of conversational agents and chat platforms, code-switching has become an integral part of written conversations in many multi-lingual communities…

2021

LIFI: Towards Linguistically Informed Frame Interpolation

ICASSP 2021accepted

Here we explore the problem of speech video interpolation. With close to 70% of web traffic, such content today forms the primary form of online communication and entertainment. Despite high performance on conventional metrics like MSE, PSNR, and SSIM, we find that the state-of-the-art frame interpo…

Cited by 0SourceScholar
2021

Meta-Learning for Low-Resource Speech Emotion Recognition

ICASSP 2021accepted

While emotion recognition is a well-studied task, it remains unexplored to a large extent in cross-lingual settings. Speech Emotion Recognition (SER) in low-resource languages poses difficulties as existing approaches for knowledge transfer do not generalize seamlessly. Probing the learning process…

Cited by 0SourceScholar
2021

Modeling financial uncertainty with multivariate temporal entropy-based curriculums

UAI 2021poster

In the financial realm, profit generation greatly relies on the complicated task of stock prediction. Lately, neural methods have shown success in exploiting stock affecting signals from textual data across news and tweets to forecast stock performance. However, the dynamic, stochastic, and variably…

Cited by 4SourcePDFScholar
2021

Multimodal Multi-Speaker Merger & Acquisition Financial Modeling: A New Task, Dataset, and Neural Baselines

ACL 2021long

Risk prediction is an essential task in financial markets. Merger and Acquisition (M&A) calls provide key insights into the claims made by company executives about the restructuring of the financial firms. Extracting vocal and textual cues from M&A calls can help model the risk associated with such…

Cited by 18SourcePDFScholar
2021

Multitask Learning for Emotionally Analyzing Sexual Abuse Disclosures

NAACL 2021long

The #MeToo movement on social media platforms initiated discussions over several facets of sexual harassment in our society. Prior work by the NLP community for automated identification of the narratives related to sexual abuse disclosures barely explored this social phenomenon as an independent tas…

2021

Quantitative Day Trading from Natural Language using Reinforcement Learning

NAACL 2021long

It is challenging to design profitable and practical trading strategies, as stock price movements are highly stochastic, and the market is heavily influenced by chaotic data across sources like news and social media. Existing NLP approaches largely treat stock prediction as a classification or regre…

2021

Stock Selection via Spatiotemporal Hypergraph Attention Network: A Learning to Rank Approach

AAAI 2021technical

Quantitative trading and investment decision making are intricate financial tasks that rely on accurate stock selection. Despite advances in deep learning that have made significant progress in the complex and highly stochastic stock prediction problem, modern solutions face two significant limitati…

Cited by 129SourcePDFScholar
2021

Suicide Ideation Detection via Social and Temporal User Representations using Hyperbolic Learning

NAACL 2021long

Recent psychological studies indicate that individuals exhibiting suicidal ideation increasingly turn to social media rather than mental health practitioners. Personally contextualizing the buildup of such ideation is critical for accurate identification of users at risk. In this work, we propose a…

Cited by 56SourcePDFScholar
2020

Augmenting NLP models using Latent Feature Interpolations

COLING 2020main

Models with a large number of parameters are prone to over-fitting and often fail to capture the underlying input distribution. We introduce Emix, a data augmentation method that uses interpolations of word embeddings and hidden layer representations to construct virtual examples. We show that Emix…

Cited by 32SourcePDFScholar
2020

GPolS: A Contextual Graph-Based Language Model for Analyzing Parliamentary Debates and Political Cohesion

COLING 2020main

Parliamentary debates present a valuable language resource for analyzing comprehensive options in electing representatives under a functional, free society. However, the esoteric nature of political speech coupled with non-linguistic aspects such as political cohesion between party members presents…

Cited by 22SourcePDFScholar
2020

Mixup Multi-Attention Multi-Tasking Model for Early-Stage Leukemia Identification

ICASSP 2020accepted

Recently, several image processing and deep learning techniques have been applied to automate the detection of Acute Lymphoblastic Leukemia cells (ALL). However, most of them have consistently focused on classification mature stage cell images into binary categories of ALL or normal cells. The real…

Cited by 0SourceScholar
2020

Mt-Gcn For Multi-Label Audio Tagging With Noisy Labels

ICASSP 2020accepted

Multi-label audio tagging is the task of predicting the types of sounds occurring in an audio clip. Recently, large-scale audio datasets such as Google's AudioSet, have allowed researchers to use deep learning techniques for this task but this comes at the cost of label noise in the datasets. Audio…

Cited by 0SourceScholar