← Search

Zheng Wu

22 accepted papers

2026

City-Scale Lane-Level Mapping From Crowdsourced Trajectories and Satellite Imagery

RA-L 2026

Lane-level maps are increasingly preferred over Standard-Definition (SD) and High-Definition (HD) maps, offering a better trade-off among detail richness, coverage breadth, and data freshness. However, constructing city-scale lane-level maps remains time-consuming and labor-intensive. To address the

Cited by 0SourceScholar
2026

Faithful Mobile GUI Agents with Guided Advantage Estimator

ICML 2026poster

Vision-language model (VLM) based graphical user interface (GUI) agents have shown strong interaction capabilities. However, they often behave unfaithfully, relying on memorized shortcuts rather than grounding actions in displayed screen evidence or user instructions. To address this, we propose **F…

Cited by 0SourceScholar
2026

GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents

AAAI 2026technical

Graphical user interface (GUI) agents have recently emerged as an intriguing paradigm for human-computer interaction, capable of automatically executing user instructions to operate intelligent terminal devices. However, when encountering out-of-distribution (OOD) instructions that violate environm

Cited by 0SourcePDFScholar
2026

See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

CVPR 2026

The advent of multimodal agents facilitates effective interaction within graphical user interface (GUI), especially in ubiquitous GUI control. However, their inability to reliably execute toggle control instructions remains a key bottleneck. To investigate this, we construct a state control benchmar

Cited by 0SourcecodeScholar
2026

Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation

CVPR 2026

The rapid development of large vision-language model (VLM) has greatly promoted the research of GUI agent. However, GUI agents still face significant challenges in handling long-horizon tasks. First, single-agent models struggle to balance high-level capabilities and low-level execution capability,

Cited by 0SourcecodeScholar
2025

Flaming-hot Initiation with Regular Execution Sampling for Large Language Models

NAACL 2025findings

Since the release of ChatGPT, large language models (LLMs) have demonstrated remarkable capabilities across various domains. A key challenge in developing these general capabilities is efficiently sourcing diverse, high-quality data. This becomes especially critical in reasoning-related tasks with s…

Cited by 2SourcePDFScholar
2025

Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents

EMNLP 2025

Graphical user interface (GUI) agents powered by multimodal large language models (MLLMs) have shown greater promise for human-interaction. However, due to the high fine-tuning cost, users often rely on open-source GUI agents or APIs offered by AI providers, which introduces a critical but underexpl

2025

OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents

ACL 2025finding

Autonomous graphical user interface (GUI) agents powered by multimodal large language models have shown great promise. However, a critical yet underexplored issue persists: over-execution, where the agent executes tasks in a fully autonomous way, without adequate assessment of its action confidence…

2025

Physics-Aware Robotic Palletization With Online Masking Inference

ICRA 2025

The efficient planning of stacking boxes, especially in the online setting where the sequence of item arrivals is unpredictable, remains a critical challenge in modern warehouse and logistics management. Existing solutions often address box size variations, but overlook their intrinsic and physical

Cited by 5SourcecodeScholar
2024

DBPF: A Framework for Efficient and Robust Dynamic Bin-Picking

RA-L 2024

Efficiency and reliability are critical in robotic bin-picking as they directly impact the productivity of automated industrial processes. However, traditional approaches, demanding static objects and fixed collisions, lead to deployment limitations, operational inefficiencies, and process unreliabi

Cited by 5SourceScholar
2024

Efficient Reinforcement Learning of Task Planners for Robotic Palletization Through Iterative Action Masking Learning

RA-L 2024

The development of robotic systems for palletization in logistics scenarios is of paramount importance, addressing critical efficiency and precision demands in supply chain management. This paper investigates the application of Reinforcement Learning (RL) in enhancing task planning for such robotic

Cited by 13SourceScholar
2023

Efficient Sim-to-real Transfer of Contact-Rich Manipulation Skills with Online Admittance Residual Learning

CoRL 2023poster

Learning contact-rich manipulation skills is essential. Such skills require the robots to interact with the environment with feasible manipulation trajectories and suitable compliance control parameters to enable safe and stable contact. However, learning these skills is challenging due to data inef…

Cited by 24SourceScholar
2023

Gradient-Based Graph Attention for Scene Text Image Super-resolution

AAAI 2023technical

Scene text image super-resolution (STISR) in the wild has been shown to be beneficial to support improved vision-based text recognition from low-resolution imagery. An intuitive way to enhance STISR performance is to explore the well-structured and repetitive layout characteristics of text and explo…

2023

Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning

ICRA 2023poster

Humans are capable of abstracting various tasks as different combinations of multiple attributes. This perspective of compositionality is vital for human rapid learning and adaption since previous experiences from related tasks can be combined to generalize across novel compositional settings. In th…

Cited by 14SourceScholar
2022

ComGAN: Unsupervised Disentanglement and Segmentation via Image Composition

NeurIPS 2022accept

We propose ComGAN, a simple unsupervised generative model, which simultaneously generates realistic images and high semantic masks under an adversarial loss and a binary regularization. In this paper, we first investigate two kinds of trivial solutions in the compositional generation process, and de…

Cited by 12SourcePDFScholar
2022

Offline-Online Learning of Deformation Model for Cable Manipulation With Graph Neural Networks

RA-L 2022

Manipulating deformable linear objects by robots has a wide range of applications, e.g., manufacturing and medical surgery. To complete such tasks, an accurate dynamics model for predicting the deformation is critical for robust control. In this letter, we deal with this challenge by proposing a hyb

Cited by 69SourceScholar
2022

Reinforcement learning with Demonstrations from Mismatched Task under Sparse Reward

CoRL 2022poster

Reinforcement learning often suffer from the sparse reward issue in real-world robotics problems. Learning from demonstration (LfD) is an effective way to eliminate this problem, which leverages collected expert data to aid online learning. Prior works often assume that the learning agent and the ex…

Cited by 6SourceScholar
2021

Learning Dense Rewards for Contact-Rich Manipulation Tasks

ICRA 2021poster

Rewards play a crucial role in reinforcement learning. To arrive at the desired policy, the design of a suitable reward function often requires significant domain expertise as well as trial-and-error. Here, we aim to minimize the effort involved in designing reward functions for contact-rich manipul…

Cited by 50SourceScholar
2020

Efficient Sampling-Based Maximum Entropy Inverse Reinforcement Learning With Application to Autonomous Driving

RA-L 2020

In the past decades, we have witnessed significant progress in the domain of autonomous driving. Advanced techniques based on optimization and reinforcement learning become increasingly powerful when solving the forward problem: given designed reward/cost functions, how we should optimize them and o

Cited by 122SourceScholar
2020

Expressing Diverse Human Driving Behavior with Probabilistic Rewards and Online Inference

IROS 2020poster

In human-robot interaction (HRI) systems, such as autonomous vehicles, understanding and representing human behavior are important. Human behavior is naturally rich and diverse. Cost/reward learning, as an efficient way to learn and represent human behavior, has been successfully applied in many dom…

Cited by 9SourceScholar
2019

Learning to Describe Scenes with Programs

ICLR 2019poster

Human scene perception goes beyond recognizing a collection of objects and their pairwise relations. We understand higher-level, abstract regularities within the scene such as symmetry and repetition. Current vision recognition modules and scene representations fall short in this dimension. In this…

Cited by 64SourcePDFScholar
2016

Differential Geometric Regularization for Supervised Learning of Classifiers

ICML 2016poster

We study the problem of supervised learning for both binary and multiclass classification from a unified geometric perspective. In particular, we propose a geometric regularization technique to find the submanifold corresponding to an estimator of the class probability P(y|\vec x). The regularizatio…

Cited by 3SourcePDFScholar