← Search

Harsh Agrawal

17 accepted papers

2026

Scaling Synthetic Task Generation for Agents via Exploration

ICLR 2026poster

Post-Training Multimodal Large Language Models (MLLMs) to build interactive agents holds promise across domains such as computer-use, web navigation, and robotics. A key challenge in scaling such post-training is lack of high-quality downstream agentic task datasets with tasks that are diverse, feas…

Cited by 0SourceScholar
2025

Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

ICLR 2025poster

Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferret-UI 2, a multimodal large language model (MLLM) designed for universal UI unde…

Cited by 0SourcePDFScholar
2025

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

CVPR 2025poster

We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end,…

Cited by 3SourcePDFScholar
2025

UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital Agents

ICCV 2025poster

We build a comprehensive online evaluation benchmark for language-conditioned multi-step task execution on mobile interfaces. Our benchmark strives to evaluate the multi-step planning, reasoning, and visual grounding capabilities of agents, using mobile user interfaces as a concrete testbed. To buil…

Cited by 0SourcePDFScholar
2024

Grounding Multimodal Large Language Models in Actions

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their associated action spaces, including both continuous and discrete actions. For conti…

Cited by 12SourcePDFScholar
2024

Large Language Models as Generalizable Policies for Embodied Tasks

ICLR 2024poster

We show that large language models (LLMs) can be adapted to be generalizable policies for embodied visual tasks. Our approach, called Large LAnguage model Reinforcement Learning Policy (LLaRP), adapts a pre-trained frozen LLM to take as input text instructions and visual egocentric observations and…

Cited by 75SourcePDFScholar
2023

Simple and Effective Synthesis of Indoor 3D Scenes

AAAI 2023technical

We study the problem of synthesizing immersive 3D indoor scenes from one or a few images. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the input images while maintaining 3D consistency. Existing approaches are highly…

2022

Housekeep: Tidying Virtual Households Using Commonsense Reasoning

ECCV 2022poster

"We introduce Housekeep, a benchmark to evaluate commonsense reasoning in the home for embodied AI. In Housekeep, an embodied agent must tidy a house by rearranging misplaced objects without explicit instructions specifying which objects need to be rearranged. Instead, the agent must learn from and…

2021

Contrast and Classify: Training Robust VQA Models

ICCV 2021poster

Recent Visual Question Answering (VQA) models have shown impressive performance on the VQA benchmark but remain sensitive to small linguistic variations in input questions. Existing approaches address this by augmenting the dataset with question paraphrases from visual question generation models or…

Cited by 35PDFcodeScholar
2021

SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation

NeurIPS 2021poster

Natural language instructions for visual navigation often use scene descriptions (e.g., bedroom) and object references (e.g., green chairs) to provide a breadcrumb trail to a goal location. This work presents a transformer-based vision-and-language navigation (VLN) agent that uses two different visu…

Cited by 64SourcePDFScholar
2021

The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal Navigation

ICCV 2021poster

It is fundamental for personal robots to reliably navigate to a specified goal. To study this task, PointGoal navigation has been introduced in simulated Embodied AI environments. Recent advances solve this PointGoal navigation task with near-perfect accuracy (99.6% success) in photo-realistically s…

Cited by 53PDFScholar
2020

Spatially Aware Multimodal Transformers for TextVQA

ECCV 2020poster

Textual cues are essential for everyday tasks like buying groceries and using public transport. To develop this assistive technology, we study the TextVQA task, i.e., reasoning about text in images to answer a question. Existing approaches are limited in their use of spatial relations and rely on fu…

2019

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

ICCV 2019poster

Diverse and accurate vision+language modeling is an important goal to retain creative freedom and maintain user engagement. However, adequately capturing the intricacies of diversity in language models is challenging. Recent works commonly resort to latent variable models augmented with more or less…

Cited by 84PDFScholar
2019

nocaps: novel object captioning at scale

ICCV 2019poster

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supe…

Cited by 420PDFcodeScholar