← Search

Yuqi Zhou

5 accepted papers

2026

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

ICML 2026oral

Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of real-world mobile usage. To this end, we introduce VenusBench-Mobile, a challenging online benchmark for evaluating general-purpose mobile GUI agents …

Cited by 0SourceScholar
2025

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents

NeurIPS 2025poster

Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object grounding and thereby achieving substantial performance gains. In this paper, we first conduct extensive analysis experime…

Cited by 0SourcecodeScholar
2025

Length-Induced Embedding Collapse in PLM-based Models

ACL 2025long

Text embeddings from PLM-based models enable a wide range of applications, yet their performance often degrades on longer texts. In this paper, we introduce a phenomenon we call Length Collapse, where embeddings of longer texts tend to cluster together. This clustering results in a distributional in…

2024

Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration

ACL 2024findings

The proliferation of Large Language Models (LLMs) has led to an influx of AI-generated content (AIGC) on the internet, transforming the corpus of Information Retrieval (IR) systems from solely human-written to a coexistence with LLM-generated content. The impact of this surge in AIGC on IR systems r…

2024

Virtual Context Enhancing Jailbreak Attacks with Special Token Injection

EMNLP 2024finding

Jailbreak attacks on large language models (LLMs) involve inducing these models to generate harmful content that violates ethics or laws, posing a significant threat to LLM security. Current jailbreak attacks face two main challenges: low success rates due to defensive measures and high resource req…

Cited by 8SourcePDFScholar