← Search

Anahita Bhiwandiwalla

6 accepted papers

2025

LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression

NAACL 2025findings

Despite recent efforts in understanding the compression impact on Large Language Models (LLMs) in terms of their downstream task performance and trustworthiness on relatively simpler uni-modal benchmarks (e.g. question answering, common sense reasoning), their detailed study on multi-modal Large Vis…

Cited by 1SourcePDFScholar
2025

Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals

NAACL 2025long

With the advent of Large Language Models (LLMs) possessing increasingly impressive capabilities, a number of Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs. Such models condition generated text on both an input image and a text prompt, enabling a variety o…

Cited by 8SourcePDFScholar
2024

SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples

CVPR 2024poster

While vision-language models (VLMs) have achieved remarkable performance improvements recently there is growing evidence that these models also posses harmful biases with respect to social attributes such as gender and race. Prior studies have primarily focused on probing such bias attributes indivi…

2024

Why do LLaVA Vision-Language Models Reply to Images in English?

EMNLP 2024finding

We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significantly increases the likelihood of the model returning an English response, regardless of the language of the query. This pa…

Cited by 4SourcePDFScholar
2023

ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning

ACL 2023long

Two-Tower Vision-Language (VL) models have shown promising improvements on various downstream VL tasks. Although the most advanced work improves performance by building bridges between encoders, it suffers from ineffective layer-by-layer utilization of uni-modal representations and cannot flexibly e…

2020

Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural Networks

ICLR 2020poster

Training with larger number of parameters while keeping fast iterations is an increasingly adopted strategy and trend for developing better performing Deep Neural Network (DNN) models. This necessitates increased memory footprint and computational requirements for training. Here we introduce a novel…

Cited by 63SourceScholar