← Search

Akash Ghosh

16 accepted papers

2026

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

IJCAI 2026

Multimodal Large Language Models(MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, particularly for multilingual and low-resource scenarios. This gap is critical in regions like rural India, where

Cited by 0Scholar
2026

CLINIC : Evaluating Multilingual Trustworthiness in Language Models for Healthcare

ICML 2026poster

Integrating language models (LMs) in healthcare systems holds great promise for improving medical workflows and decision-making. However, a critical barrier to their real-world adoption is the lack of reliable evaluation of their trustworthiness, especially in multilingual healthcare settings. Exist…

Cited by 0SourceScholar
2026

The Cylindrical Representation Hypothesis for Language Model Steering

ICML 2026poster

Steering is a widely used technique for controlling large language models, yet its effects are often unstable and hard to predict. Existing theoretical accounts are largely based on the Linear Representation Hypothesis (LRH). While LRH assumes that concepts can be orthogonalized for lossless control…

Cited by 0SourceScholar
2025

A Survey of Multilingual Reasoning in Language Models

EMNLP 2025

While reasoning and multilingual capabilities in Language Models (LMs) have achieved remarkable progress in recent years, their integration into a unified paradigm—multilingual reasoning—is at a nascent stage. Multilingual reasoning requires language models to handle logical reasoning across languag

2025

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Culture

EMNLP 2025

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchmarks with a generic or global scope, DRISHTIKON offers deep, fine-grained coverag

Cited by 0SourcePDFScholar
2025

Infogen: Generating Complex Statistical Infographics from Documents

ACL 2025long

Statistical infographics are powerful tools that simplify complex data into visually engaging and easy-to-understand formats. Despite advancements in AI, particularly with LLMs, existing efforts have been limited to generating simple charts, with no prior work addressing the creation of complex info…

Cited by 0SourcePDFScholar
2025

Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sports

EMNLP 2025

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce CultSportQA , a benchmark designed to assess LMs’ understanding of traditional sports across 60 countries and 6 continents, encom

Cited by 0SourcePDFScholar
2025

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

EMNLP 2025

With the increasing use of Retrieval-Augmented Generation (RAG), strong retrieval models have become more important than ever. In healthcare, multimodal retrieval models that combine information from both text and images offer major advantages for many downstream tasks such as question answering, cr

2025

RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples

EMNLP 2025

Reward models are essential for aligning large language models (LLMs) with human preferences. However, most open-source multilingual reward models are primarily trained on preference datasets in high-resource languages, resulting in unreliable reward signals for low-resource Indic languages. Collect

Cited by 0SourcePDFScholar
2025

SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models’ Knowledge of Indian Culture

ACL 2025finding

Language models (LMs) are indispensable tools shaping modern workflows, but their global effectiveness depends on understanding local socio-cultural contexts. To address this, we introduce SANSKRITI, a benchmark designed to evaluate language models’ comprehension of India’s rich cultural diversity.…

Cited by 0SourcePDFScholar
2024

A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

EMNLP 2024finding

The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks. However, the proliferation of FMs brings forth a critical challenge: the potential to generate hallucinated outputs, particularly in high-stakes appli…

Cited by 17SourcePDFScholar
2024

CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare

AAAI 2024technical

In the era of modern healthcare, swiftly generating medical question summaries is crucial for informed and timely patient care. Despite the increasing complexity and volume of medical data, existing studies have focused solely on text-based summarization, neglecting the integration of visual informa…

2024

From Sights to Insights: Towards Summarization of Multimodal Clinical Documents

ACL 2024long

The advancement of Artificial Intelligence is pivotal in reshaping healthcare, enhancing diagnostic precision, and facilitating personalized treatment strategies. One major challenge for healthcare professionals is quickly navigating through long clinical documents to provide timely and effective so…

2024

HealthAlignSumm : Utilizing Alignment for Multimodal Summarization of Code-Mixed Healthcare Dialogues

EMNLP 2024finding

As generative AI progresses, collaboration be-tween doctors and AI scientists is leading to thedevelopment of personalized models to stream-line healthcare tasks and improve productivity.Summarizing doctor-patient dialogues has be-come important, helping doctors understandconversations faster and im…

2024

How Robust Are the QA Models for Hybrid Scientific Tabular Data? A Study Using Customized Dataset

COLING 2024main

Question-answering (QA) on hybrid scientific tabular and textual data deals with scientific information, and relies on complex numerical reasoning. In recent years, while tabular QA has seen rapid progress, understanding their robustness on scientific information is lacking due to absence of any ben…

Cited by 2SourcePDFScholar
2024

Parameter-Efficient Instruction Tuning of Large Language Models For Extreme Financial Numeral Labelling

NAACL 2024long

We study the problem of automatically annotating relevant numerals (GAAP metrics) occurring in the financial documents with their corresponding XBRL tags. Different from prior works, we investigate the feasibility of solving this extreme classification problem using a generative paradigm through ins…