← Search

Sriparna Saha

29 accepted papers

2026

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

IJCAI 2026

Multimodal Large Language Models(MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, particularly for multilingual and low-resource scenarios. This gap is critical in regions like rural India, where

Cited by 0Scholar
2026

CLINIC : Evaluating Multilingual Trustworthiness in Language Models for Healthcare

ICML 2026poster

Integrating language models (LMs) in healthcare systems holds great promise for improving medical workflows and decision-making. However, a critical barrier to their real-world adoption is the lack of reliable evaluation of their trustworthiness, especially in multilingual healthcare settings. Exist…

Cited by 0SourceScholar
2026

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

AAAI 2026technical

Indian poetry, known for its linguistic complexity and deep cultural resonance, has a rich and varied heritage spanning thousands of years. However, its layered meanings, cultural allusions, and sophisticated grammatical constructions often pose challenges for comprehension, especially for non-nativ

Cited by 0SourcePDFScholar
2026

Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances

AAAI 2026technical

Existing approaches to complaint analysis largely rely on unimodal, short-form content such as tweets or product reviews. This work advances the field by leveraging multimodal, multi-turn customer support dialogues—where users often share both textual complaints and visual evidence (e.g., screenshot

Cited by 0SourcePDFScholar
2025

A Survey of Multilingual Reasoning in Language Models

EMNLP 2025

While reasoning and multilingual capabilities in Language Models (LMs) have achieved remarkable progress in recent years, their integration into a unified paradigm—multilingual reasoning—is at a nascent stage. Multilingual reasoning requires language models to handle logical reasoning across languag

2025

COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation

ACL 2025long

Despite progress in comment-aware multimodal and multilingual summarization for English and Chinese, research in Indian languages remains limited. This study addresses this gap by introducing COSMMIC, a pioneering comment-sensitive multimodal, multilingual dataset featuring nine major Indian languag…

2025

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Culture

EMNLP 2025

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchmarks with a generic or global scope, DRISHTIKON offers deep, fine-grained coverag

Cited by 0SourcePDFScholar
2025

Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation

EMNLP 2025

Poetry is an expressive form of art that invites multiple interpretations, as readers often bring their own emotions, experiences, and cultural backgrounds into their understanding of a poem. Recognizing this, we aim to generate images for poems and improve these images in a zero-shot setting, enabl

Cited by 0SourcePDFScholar
2025

Infogen: Generating Complex Statistical Infographics from Documents

ACL 2025long

Statistical infographics are powerful tools that simplify complex data into visually engaging and easy-to-understand formats. Despite advancements in AI, particularly with LLMs, existing efforts have been limited to generating simple charts, with no prior work addressing the creation of complex info…

Cited by 0SourcePDFScholar
2025

Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sports

EMNLP 2025

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce CultSportQA , a benchmark designed to assess LMs’ understanding of traditional sports across 60 countries and 6 continents, encom

Cited by 0SourcePDFScholar
2025

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

EMNLP 2025

With the increasing use of Retrieval-Augmented Generation (RAG), strong retrieval models have become more important than ever. In healthcare, multimodal retrieval models that combine information from both text and images offer major advantages for many downstream tasks such as question answering, cr

2025

Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models

COLING 2025main

The task of text-to-image generation has encountered significant challenges when applied to literary works, especially poetry. Poems are a distinct form of literature, with meanings that frequently transcend beyond the literal words. To address this shortcoming, we propose a PoemToPixel framework de…

2025

RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples

EMNLP 2025

Reward models are essential for aligning large language models (LLMs) with human preferences. However, most open-source multilingual reward models are primarily trained on preference datasets in high-resource languages, resulting in unreliable reward signals for low-resource Indic languages. Collect

Cited by 0SourcePDFScholar
2025

SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models’ Knowledge of Indian Culture

ACL 2025finding

Language models (LMs) are indispensable tools shaping modern workflows, but their global effectiveness depends on understanding local socio-cultural contexts. To address this, we introduce SANSKRITI, a benchmark designed to evaluate language models’ comprehension of India’s rich cultural diversity.…

Cited by 0SourcePDFScholar
2024

A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

EMNLP 2024finding

The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks. However, the proliferation of FMs brings forth a critical challenge: the potential to generate hallucinated outputs, particularly in high-stakes appli…

Cited by 17SourcePDFScholar
2024

Action and Reaction Go Hand in Hand! a Multi-modal Dialogue Act Aided Sarcasm Identification

COLING 2024main

Sarcasm primarily involves saying something but “meaning the opposite” or “meaning something completely different” in order to convey a particular tone or mood. In both the above cases, the “meaning” is reflected by the communicative intention of the speaker, known as dialogue acts. In this paper, w…

2024

CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare

AAAI 2024technical

In the era of modern healthcare, swiftly generating medical question summaries is crucial for informed and timely patient care. Despite the increasing complexity and volume of medical data, existing studies have focused solely on text-based summarization, neglecting the integration of visual informa…

2024

Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development

ACL 2024findings

The mining of adverse drug events (ADEs) is pivotal in pharmacovigilance, enhancing patient safety by identifying potential risks associated with medications, facilitating early detection of adverse events, and guiding regulatory decision-making. Traditional ADE detection methods are reliable but sl…

2024

From Sights to Insights: Towards Summarization of Multimodal Clinical Documents

ACL 2024long

The advancement of Artificial Intelligence is pivotal in reshaping healthcare, enhancing diagnostic precision, and facilitating personalized treatment strategies. One major challenge for healthcare professionals is quickly navigating through long clinical documents to provide timely and effective so…

2024

HealthAlignSumm : Utilizing Alignment for Multimodal Summarization of Code-Mixed Healthcare Dialogues

EMNLP 2024finding

As generative AI progresses, collaboration be-tween doctors and AI scientists is leading to thedevelopment of personalized models to stream-line healthcare tasks and improve productivity.Summarizing doctor-patient dialogues has be-come important, helping doctors understandconversations faster and im…

2024

MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention

ACL 2024long

In the digital world, memes present a unique challenge for content moderation due to their potential to spread harmful content. Although detection methods have improved, proactive solutions such as intervention are still limited, with current research focusing mostly on text-based content, neglectin…

2024

Seeing Is Believing! towards Knowledge-Infused Multi-modal Medical Dialogue Generation

COLING 2024main

Over the last few years, artificial intelligence-based clinical assistance has gained immense popularity and demand in telemedicine, including automatic disease diagnosis. Patients often describe their signs and symptoms to doctors using visual aids, which provide vital evidence for identifying a me…

2024

ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos

ACL 2024findings

In an era of rapidly evolving internet technology, the surge in multimodal content, including videos, has expanded the horizons of online communication. However, the detection of toxic content in this diverse landscape, particularly in low-resource code-mixed languages, remains a critical challenge.…

2023

Can you Summarize my learnings? Towards Perspective-based Educational Dialogue Summarization

EMNLP 2023long findings

The steady increase in the utilization of Virtual Tutors (VT) over recent years has allowed for a more efficient, personalized, and interactive AI-based learning experiences. A vital aspect in these educational chatbots is summarizing the conversations between the VT and the students, as it is criti…

Cited by 0SourceScholar
2023

Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models

EMNLP 2023long main

Temporal reasoning represents a vital component of human communication and understanding, yet remains an underexplored area within the context of Large Language Models (LLMs). Despite LLMs demonstrating significant proficiency in a range of tasks, a comprehensive, large-scale analysis of their tempo…

Cited by 0SourceScholar
2023

Peeking inside the black box: A Commonsense-aware Generative Framework for Explainable Complaint Detection

ACL 2023long

Complaining is an illocutionary act in which the speaker communicates his/her dissatisfaction with a set of circumstances and holds the hearer (the complainee) answerable, directly or indirectly. Considering breakthroughs in machine learning approaches, the complaint detection task has piqued the in…

2022

A Shoulder to Cry on: Towards A Motivational Virtual Assistant for Assuaging Mental Agony

NAACL 2022long

Mental Health Disorders continue plaguing humans worldwide. Aggravating this situation is the severe shortage of qualified and competent mental health professionals (MHPs), which underlines the need for developing Virtual Assistants (VAs) that can assist MHPs. The data+ML for automation can come fro…

Cited by 27SourcePDFScholar
2022

Many Hands Make Light Work: Using Essay Traits to Automatically Score Essays

NAACL 2022long

Most research in the area of automatic essay grading (AEG) is geared towards scoring the essay holistically while there has also been little work done on scoring individual essay traits. In this paper, we describe a way to score essays using a multi-task learning (MTL) approach, where scoring the es…

2021

Towards Sentiment and Emotion aided Multi-modal Speech Act Classification in Twitter

NAACL 2021long

Speech Act Classification determining the communicative intent of an utterance has been investigated widely over the years as a standalone task. This holds true for discussion in any fora including social media platform such as Twitter. But the emotional state of the tweeter which has a considerable…

Cited by 36SourcePDFScholar