← Search

Xinbin Yuan

5 accepted papers

2026

LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI Agent

ICLR 2026poster

Although agents based on multimodal large language models (MLLMs) demonstrate proficiency in general short-term graphical user interface (GUI) tasks, their robustness remains a significant challenge for handling complex long-horizon tasks in dynamic environments . In response, the LongHorizonUI fram…

Cited by 0SourcecodeScholar
2026

Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models

CVPR 2026

While Multimodal Large Language Models (MLLMs) excel at vision-language tasks, the cost of their language-driven training on internal visual foundational competence remains unclear. In this paper, we conduct a detailed diagnostic analysis to unveil a pervasive issue: visual representation degradatio

Cited by 0SourceScholar
2026

Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection

AAAI 2026technical

In this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-C

Cited by 0SourcePDFScholar
2025

Photography Perspective Composition: Towards Aesthetic Perspective Recommendation

NeurIPS 2025poster

Traditional photography composition approaches are dominated by 2D cropping-based methods. However, these methods fall short when scenes contain poorly arranged subjects. Professional photographers often employ perspective adjustment as a form of 3D recomposition, modifying the projected 2D relation…

Cited by 0SourceScholar
2025

SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

NeurIPS 2025poster

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to precise interface elements remains challenging—especially in complex, high-resolution, professional environments. Tradit…

Cited by 0SourceScholar