← Search

Jinyuan Li

5 accepted papers

2026

Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution

ICML 2026poster

Multimodal Large Language Model (MLLM) agents facilitate Graphical User Interface (GUI) automation but struggle with long-horizon, cross-application tasks due to limited context windows. While memory systems provide a viable solution, existing paradigms struggle to adapt to dynamic GUI environments,…

Cited by 0SourceScholar
2026

Training Data Efficiency in Multimodal Process Reward Models

ICML 2026poster

Multimodal Process Reward Models (MPRMs) are central to step-level supervision for visual reasoning in MLLMs. Training MPRMs typically requires large-scale Monte Carlo (MC)-annotated corpora, incurring substantial training cost. This paper studies the data efficiency for MPRM training. Our prelimina…

Cited by 0SourceScholar
2025

VP-MEL: Visual Prompts Guided Multimodal Entity Linking

ACL 2025finding

Multimodal entity linking (MEL), a task aimed at linking mentions within multimodal contexts to their corresponding entities in a knowledge base (KB), has attracted much attention due to its wide applications in recent years. However, existing MEL methods often rely on mention words as retrieval cue…

Cited by 0SourcePDFScholar
2024

LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition

ACL 2024findings

Grounded Multimodal Named Entity Recognition (GMNER) is a nascent multimodal task that aims to identify named entities, entity types and their corresponding visual regions. GMNER task exhibits two challenging properties: 1) The weak correlation between image-text pairs in social media results in a s…

2023

Prompting ChatGPT in MNER: Enhanced Multimodal Named Entity Recognition with Auxiliary Refined Knowledge

EMNLP 2023long findings

Multimodal Named Entity Recognition (MNER) on social media aims to enhance textual entity prediction by incorporating image-based clues. Existing studies mainly focus on maximizing the utilization of pertinent image information or incorporating external knowledge from explicit knowledge bases. Howev…

Cited by 0SourcecodeScholar