← Search

Atharva Kulkarni

9 accepted papers

2026

Disentangling Geometry, Performance, and Training in Language Models

ICML 2026spotlight

Geometric properties of Transformer weights, particularly the unembedding matrix, have been widely useful in language model interpretability research. Yet, their utility for estimating downstream performance remains unclear. In this work, we systematically investigate the relationship between model …

Cited by 0SourceScholar
2025

Evaluating Evaluation Metrics – The Mirage of Hallucination Detection

EMNLP 2025

Hallucinations pose a significant obstacle to the reliability and widespread adoption of language models, yet their accurate measurement remains a persistent challenge. While many task- and domain-specific metrics have been proposed to assess faithfulness and factuality concerns, the robustness and

Cited by 0SourcePDFScholar
2025

Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models

NAACL 2025findings

The advent of Music-Language Models has greatly enhanced the automatic music generation capability of AI systems, but they are also limited in their coverage of the musical genres and cultures of the world. We present a study of the datasets and research papers for music generation and quantify the…

2024

Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis

EMNLP 2024main

In this study, we introduce ANGST, a novel, first of its kind benchmark for depression-anxiety comorbidity classification from social media posts. Unlike contemporary datasets that often oversimplify the intricate interplay between different mental health disorders by treating them as isolated condi…

2023

Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model

EMNLP 2023long main

Large language models (LLMs) have recently reached an impressive level of linguistic capability, prompting comparisons with human language skills. However, there have been relatively few systematic inquiries into the linguistic capabilities of the latest generation of LLMs, and those studies that do…

Cited by 0SourceScholar
2023

Learning and Reasoning Multifaceted and Longitudinal Data for Poverty Estimates and Livelihood Capabilities of Lagged Regions in Rural India

IJCAI 2023poster

Poverty is a multifaceted phenomenon linked to the lack of capabilities of households to earn a sustainable livelihood, increasingly being assessed using multidimensional indicators. Its spatial pattern depends on social, economic, political, and regional variables. Artificial intelligence has shown…

Cited by 2SourcePDFScholar
2023

The student becomes the master: Outperforming GPT3 on Scientific Factual Error Correction

EMNLP 2023long findings

Due to the prohibitively high cost of creating error correction datasets, most Factual Claim Correction methods rely on a powerful verification model to guide the correction process. This leads to a significant drop in performance in domains like Scientific Claim Correction, where good verification…

Cited by 0SourceScholar
2022

Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter

EMNLP 2022main

The widespread diffusion of medical and political claims in the wake of COVID-19 has led to a voluminous rise in misinformation and fake news. The current vogue is to employ manual fact-checkers to efficiently classify and verify such data to combat this avalanche of claim-ridden misinformation. How…

2022

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

ACL 2022long

Indirect speech such as sarcasm achieves a constellation of discourse goals in human communication. While the indirectness of figurative language warrants speakers to achieve certain pragmatic goals, it is challenging for AI agents to comprehend such idiosyncrasies of human communication. Though sar…