← Search

Pranshu Pandya

3 accepted papers

2025

NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

NAACL 2025findings

Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. Due to extensive training on vast amounts of human-curated data, large language models (LLMs) and vision languag…

Cited by 2SourcePDFScholar
2024

Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets

EMNLP 2024main

Language models, characterized by their black-box nature, often hallucinate and display sensitivity to input perturbations, causing concerns about trust. To enhance trust, it is imperative to gain a comprehensive understanding of the model’s failure modes and develop effective strategies to improve…

Cited by 1SourcePDFScholar
2024

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

ACL 2024findings

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities of visual question-answering multimodal language models in reasoning with flowch…

Cited by 11SourcePDFScholar