← Search

Yinheng Li

4 accepted papers

2025

Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems

NeurIPS 2025poster

Transformers and their attention mechanism have been revolutionary in the field of Machine Learning. While originally proposed for the language data, they quickly found their way to the image, video, graph, etc. data modalities with various signal geometries. Despite this versatility, generalizing t…

Cited by 0SourceScholar
2025

VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

ICLR 2025poster

Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text or static imagery can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focus- ing on text or static image inputs. To bridge this ga…

Cited by 3SourcePDFScholar
2025

WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models

ACL 2025short

Graphical User Interface (GUI) automation relies on accurate GUI grounding. However, obtaining large-scale, high-quality labeled data remains a key challenge, particularly in desktop environments like Windows Operating System (OS). Existing datasets primarily focus on structured web-based elements,…

2025

Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

ICML 2025poster

Large language models (LLMs) show potential as computer agents, enhancing productivity and software accessibility in multi-modal tasks. However, measuring agent performance in sufficiently realistic and complex environments becomes increasingly challenging as: (i) most benchmarks are limited to sp…