← Search

Vivek Gupta

34 accepted papers

2025

Enhancing Temporal Understanding in LLMs for Semi-structured Tables

NAACL 2025findings

Temporal reasoning over tabular data presents substantial challenges for large language models (LLMs), as evidenced by recent research. In this study, we conduct a comprehensive analysis of temporal datasets to pinpoint the specific limitations of LLMs. Our investigation leads to enhancements in Tem…

Cited by 3SourcePDFScholar
2025

Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents

EMNLP 2025

Flowcharts are a critical tool for visualizing decision-making processes. However, their non-linear structure and complex visual-textual relationships make it challenging to interpret them using LLMs, as vision-language models frequently hallucinate nonexistent connections and decision paths when an

2025

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

ACL 2025long

Publicly significant images from events carry valuable contextual information with applications in domains such as journalism and education. However, existing methodologies often struggle to accurately extract this contextual relevance from images. To address this challenge, we introduce GETREASON (…

Cited by 0SourcePDFScholar
2025

H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables

NAACL 2025long

Tabular reasoning involves interpreting natural language queries about tabular data, which presents a unique challenge of combining language understanding with structured data analysis. Existing methods employ either textual reasoning, which excels in semantic interpretation but struggles with mathe…

2025

LLM-Symbolic Integration for Robust Temporal Tabular Reasoning

ACL 2025finding

Temporal tabular question answering presents a significant challenge for Large Language Models (LLMs), requiring robust reasoning over structured data—a task where traditional prompting methods often fall short. These methods face challenges such as memorization, sensitivity to table size, and reduc…

2025

Leveraging LLM For Synchronizing Information Across Multilingual Tables

NAACL 2025long

The vast amount of online information today poses challenges for non-English speakers, as much of it is concentrated in high-resource languages such as English and French. Wikipedia reflects this imbalance, with content in low-resource languages frequently outdated or incomplete. Recent research has…

Cited by 0SourcePDFScholar
2025

M-Help: Using Social Media Data to Detect Mental Health Help-Seeking Signals

EMNLP 2025

Mental health disorders are a global crisis. While various datasets exist for detecting such disorders, there remains a critical gap in identifying individuals actively seeking help. This paper introduces a novel dataset, M-Help, specifically designed to detect help-seeking behavior on social media.

Cited by 0SourcePDFScholar
2025

MAPWise: Evaluating Vision-Language Models for Advanced Map Queries

NAACL 2025long

Vision-language models (VLMs) excel at tasks requiring joint understanding of visual and linguistic information. A particularly promising yet under-explored application for these models lies in answering questions based on various kinds of maps. This study investigates the efficacy of VLMs in answer…

2025

NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

NAACL 2025findings

Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. Due to extensive training on vast amounts of human-curated data, large language models (LLMs) and vision languag…

Cited by 2SourcePDFScholar
2025

TABARD: A Novel Benchmark for Tabular Anomaly Analysis, Reasoning and Detection

EMNLP 2025

We study the capabilities of large language models (LLMs) in detecting fine-grained anomalies in tabular data. Specifically, we examine: (1) how well LLMs can identify diverse anomaly types including factual, logical, temporal, and value-based errors; (2) the impact of prompt design and prompting st

Cited by 0SourcePDFScholar
2025

TRANSIENTTABLES: Evaluating LLMs’ Reasoning on Temporally Evolving Semi-structured Tables

NAACL 2025long

Humans continuously make new discoveries, and understanding temporal sequence of events leading to these breakthroughs is essential for advancing science and society. This ability to reason over time allows us to identify future steps and understand the effects of financial and political decisions o…

2025

TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation

ACL 2025finding

Evaluating tables qualitatively and quantitatively poses a significant challenge, as standard metrics often overlook subtle structural and content-level discrepancies. To address this, we propose a rubric-based evaluation framework that integrates multi-level structural descriptors with fine-grained…

2025

Weaver: Interweaving SQL and LLM for Table Reasoning

EMNLP 2025

Querying tables with unstructured data is challenging due to the presence of text (or image), either embedded in the table or in external paragraphs, which traditional SQL struggles to process, especially for tasks requiring semantic reasoning. While Large Language Models (LLMs) excel at understandi

2024

ChartCheck: Explainable Fact-Checking over Real-World Chart Images

ACL 2024findings

Whilst fact verification has attracted substantial interest in the natural language processing community, verifying misinforming statements against data visualizations such as charts has so far been overlooked. Charts are commonly used in the real-world to summarize and com municate key information,…

2024

Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets

EMNLP 2024main

Language models, characterized by their black-box nature, often hallucinate and display sensitivity to input perturbations, causing concerns about trust. To enhance trust, it is imperative to gain a comprehensive understanding of the model’s failure modes and develop effective strategies to improve…

Cited by 1SourcePDFScholar
2024

Evaluating LLMs’ Mathematical Reasoning in Financial Document Question Answering

ACL 2024findings

Large Language Models (LLMs), excel in natural language understanding, but their capability for complex mathematical reasoning with a hybrid of structured tables and unstructured text remain uncertain. This study explores LLMs’ mathematical reasoning on four financial tabular question-answering data…

Cited by 25SourcePDFScholar
2024

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

ACL 2024findings

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities of visual question-answering multimodal language models in reasoning with flowch…

Cited by 11SourcePDFScholar
2024

Knowledge-Aware Reasoning over Multimodal Semi-structured Tables

EMNLP 2024finding

Existing datasets for tabular question answering typically focus exclusively on text within cells. However, real-world data is inherently multimodal, often blending images such as symbols, faces, icons, patterns, and charts with textual content in tables. With the evolution of AI models capable of m…

Cited by 3SourcePDFScholar
2024

Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness

EMNLP 2024finding

Chart question answering (CQA) is a crucial area of Visual Language Understanding. However, the robustness and consistency of current Visual Language Models (VLMs) in this field remain under-explored. This paper evaluates state-of-the-art VLMs on comprehensive datasets, developed specifically for th…

Cited by 4SourcePDFScholar
2023

Exploring the Numerical Reasoning Capabilities of Language Models: A Comprehensive Analysis on Tabular Data

EMNLP 2023long findings

Numerical data plays a crucial role in various real-world domains like finance, economics, and science. Thus, understanding and reasoning with numbers are essential in these fields. Recent benchmarks have assessed the numerical reasoning abilities of language models, revealing their limitations in l…

Cited by 0SourceScholar
2023

InfoSync: Information Synchronization across Multilingual Semi-structured Tables

ACL 2023findings

Information Synchronization of semi-structured data across languages is challenging. For example, Wikipedia tables in one language need to be synchronized with others. To address this problem, we introduce a new dataset InfoSync and a two-step method for tabular synchronization. InfoSync contains 10…

Cited by 4SourcePDFScholar
2023

MANER: Multi-Agent Neural Rearrangement Planning of Objects in Cluttered Environments

RA-L 2023

Object rearrangement is a fundamental problem in robotics with various practical applications ranging from managing warehouses to cleaning and organizing home kitchens. While existing research has primarily focused on single-agent solutions, real-world scenarios often require multiple robots to work

Cited by 2SourceScholar
2023

TempTabQA: Temporal Question Answering for Semi-Structured Tables

EMNLP 2023long main

Semi-structured data, such as Infobox tables, often include temporal information about entities, either implicitly or explicitly. Can current NLP systems reason about such information in semi-structured tables? To tackle this question, we introduce the task of temporal question answering on semi-str…

Cited by 0SourceScholar
2022

Bilingual Tabular Inference: A Case Study on Indic Languages

NAACL 2022long

Existing research on Tabular Natural Language Inference (TNLI) exclusively examines the task in a monolingual setting where the tabular premise and hypothesis are in the same language. However, due to the uneven distribution of text resources on the web across languages, it is common to have the tab…

Cited by 1SourcePDFScholar
2022

IndicXNLI: Evaluating Multilingual Inference for Indian Languages

EMNLP 2022main

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce INDICXNLI, an NLI dataset for 11 Indic languages. It has been created by high-quality machine translation o…

2022

Leveraging Data Recasting to Enhance Tabular Reasoning

EMNLP 2022finding

Creating challenging tabular inference data is essential for learning complex reasoning. Prior work has mostly relied on two data generation strategies. The first is human annotation, which yields linguistically diverse data but is difficult to scale. The second category for creation is synthetic ge…

Cited by 7SourcePDFScholar
2022

Realistic Data Augmentation Framework for Enhancing Tabular Reasoning

EMNLP 2022finding

Existing approaches to constructing training data for Natural Language Inference (NLI) tasks, such as for semi-structured table reasoning, are either via crowdsourcing or fully automatic methods. However, the former is expensive and time consuming and thus limits scale, and the latter often produces…

Cited by 4SourcePDFScholar
2022

Right for the Right Reason: Evidence Extraction for Trustworthy Tabular Reasoning

ACL 2022long

When pre-trained contextualized embedding-based models developed for unstructured data are adapted for structured tabular data, they perform admirably. However, recent probing studies show that these models use spurious correlations, and often predict inference labels by focusing on false evidence o…

Cited by 13SourcePDFScholar
2021

TabPert : An Effective Platform for Tabular Perturbation

EMNLP 2021system demonstrations

To grasp the true reasoning ability, the Natural Language Inference model should be evaluated on counterfactual data. TabPert facilitates this by generation of such counterfactual data for assessing model tabular reasoning issues. TabPert allows the user to update a table, change the hypothesis, cha…