← Search

Danielle Bitterman

9 accepted papers

2025

KScope: A Framework for Characterizing the Knowledge Status of Language Models

NeurIPS 2025poster

Characterizing a large language model's (LLM's) knowledge of a given question is challenging. As a result, prior work has primarily examined LLM behavior under knowledge conflicts, where the model's internal parametric memory contradicts information in the external context. However, this does not fu…

Cited by 0SourceScholar
2025

Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

AAAI 2025technical

Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report su…

Cited by 0SourcePDFScholar
2025

Sparse Autoencoder Features for Classifications and Transferability

EMNLP 2025

Sparse Autoencoders (SAEs) provide potential for uncovering structured, human-interpretable representations in Large Language Models (LLMs), making them a crucial tool for transparent and controllable AI systems. We systematically analyze SAE for interpretable feature extraction from LLMs in safety-

2025

When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy

EMNLP 2025

Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, the extent to which LRMs can think in other languages is less studied. This is as important as answer accuracy for real-world applications since users may find the thinking tr

2025

WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation

NAACL 2025findings

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets derived from national medical examinations have long served as…

2024

Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias

NeurIPS 2024poster

Large language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data. In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessin…

2024

Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks

EMNLP 2024finding

Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. This is particularly crucial for drug names, where patients often use brand names like Advil or Tylenol instead of their generic equivalents. To st…

2024

When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?

EMNLP 2024finding

The introduction of Large Language Models (LLMs) has advanced data representation and analysis, bringing significant progress in their use for medical questions and answering. Despite these advancements, integrating tabular data, especially numerical data pivotal in clinical contexts, into LLM parad…

Cited by 7SourcePDFScholar
2023

Measuring Pointwise $\mathcal{V}$-Usable Information In-Context-ly

EMNLP 2023long findings

In-context learning (ICL) is a new learning paradigm that has gained popularity along with the development of large language models. In this work, we adapt a recently proposed hardness metric, pointwise $\mathcal{V}$-usable information (PVI), to an in-context version (in-context PVI). Compared to th…

Cited by 0SourcecodeScholar