← Search

Gaurav Verma

10 accepted papers

2025

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations

ACL 2025long

State-of-the-art multimodal web agents, powered by Multimodal Large Language Models (MLLMs), can autonomously execute many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). Current strategies for building web agents rely on (i) the generalizability of u…

Cited by 0SourcePDFScholar
2025

Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use

NAACL 2025long

Adverse Drug Reactions (ADRs) from psychiatric medications are the leading cause of hospitalizations among mental health patients. With healthcare systems and online communities facing limitations in resolving ADR-related issues, Large Language Models (LLMs) have the potential to fill this gap. Desp…

2024

A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech

ACL 2024long

Violence-provoking speech – speech that implicitly or explicitly promotes violence against the members of the targeted community, contributed to a massive surge in anti-Asian crimes during the COVID-19 pandemic. While previous works have characterized and built tools for detecting other forms of har…

Cited by 1SourcePDFScholar
2024

Cross-Modal Projection in Multimodal LLMs Doesn’t Really Project Visual Attributes to Textual Space

ACL 2024short

Multimodal large language models (MLLMs) like LLaVA and GPT-4(V) enable general-purpose conversations about images with the language modality. As off-the-shelf MLLMs may have limited capabilities on images from domains like dermatology and agriculture, they must be fine-tuned to unlock domain-specif…

2024

MM-SOC: Benchmarking Multimodal Large Language Models in Social Media Platforms

ACL 2024findings

Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces. Multimodal Large Language Models (MLLMs) have emerged as a promising…

2023

Adversarial Robustness of Prompt-based Few-Shot Learning for Natural Language Understanding

ACL 2023findings

State-of-the-art few-shot learning (FSL) methods leverage prompt-based fine-tuning to obtain remarkable results for natural language understanding (NLU) tasks. While much of the prior FSL methods focus on improving downstream task performance, there is a limited understanding of the adversarial robu…

2023

Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning

ACL 2023long

The robustness of multimodal deep learning models to realistic changes in the input text is critical for applicability on important tasks such as text-to-image retrieval and cross-modal entailment. To measure robustness, several existing approaches edit the text data, but without leveraging the cros…

2023

Learning the Visualness of Text Using Large Vision-Language Models

EMNLP 2023long main

Visual text evokes an image in a person's mind, while non-visual text fails to do so. A method to automatically detect visualness in text will enable text-to-image retrieval and generation models to augment text with relevant images. This is particularly challenging with long-form text as text-to-im…

Cited by 0SourceScholar
2022

Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions

EMNLP 2022main

As multimodal learning finds applications in a wide variety of high-stakes societal tasks, investigating their robustness becomes important. Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks. In this work, we invest…

Cited by 8SourcePDFScholar