← Search

Shrinidhi Kumbhar

3 accepted papers

2026

Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding

CVPR 2026

Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusion vision-language models (DVLMs) have shown strong performance in multimodal reasoning, offering bidirectional attention

Cited by 0SourceScholar
2025

ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints

ICLR 2025poster

Reasoning about Actions and Change (RAC) has historically played a pivotal role in solving foundational AI problems, such as the frame problem. It has driven advancements in AI fields, such as non-monotonic and commonsense reasoning. RAC remains crucial for AI systems that operate in dynamic environ…

Cited by 2SourcePDFScholar
2025

Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents

NAACL 2025findings

Materials discovery and design are essential for advancing technology across various industries by enabling the development of application-specific materials. Recent research has leveraged Large Language Models (LLMs) to accelerate this process. We explore the potential of LLMs to generate viable hy…