← Search

Debjyoti Mondal

4 accepted papers

2025

From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding

ACL 2025finding

Accurately grounding visual and textual elements within mobile user interfaces (UIs) remains a significant challenge for Vision-Language Models (VLMs). Visual grounding, a critical task in this domain, involves identifying the most relevant UI element or region based on a natural language query—a pr…

2025

RG-VQA: Leveraging Retriever-Generator Pipelines for Knowledge Intensive Visual Question Answering

EMNLP 2025

In this paper, we propose a method to improve the reasoning capabilities of Visual Question Answering (VQA) systems by integrating Dense Passage Retrievers (DPRs) with Vision Language Models (VLMs). While recent works focus on the application of knowledge graphs and chain-of-thought reasoning, we re

2024

KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning

AAAI 2024technical

Large Language Models (LLMs) have demonstrated impressive performance in natural language processing tasks by leveraging chain of thought (CoT) that enables step-by-step thinking. Extending LLMs with multimodal capabilities is the recent interest, but incurs computational cost and requires substanti…

Cited by 41SourcePDFScholar