← Search

Zhihao Wu

32 accepted papers

2026

AgentSwift: Efficient LLM Agent Design via Value-Guided Hierarchical Search

AAAI 2026technical

Large language model (LLM) agents have demonstrated strong capabilities across diverse domains, yet automated agent design remains a significant challenge. Current automated agent design approaches are often constrained by limited search spaces that primarily optimize workflows but fail to integrate

Cited by 0SourcePDFScholar
2026

Beyond Local Patterns: Multiscale Inconsistency Learning for Graph Anomaly Detection

AAAI 2026technical

Graph anomaly detection is emerging as a critical technology for addressing increasingly complex and dynamic risk environments. Although unsupervised graph anomaly detection has advanced under the graph representation learning, directly applying these paradigms remains fundamentally misaligned with

Cited by 0SourcePDFScholar
2026

Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis

AAAI 2026technical

Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention m

Cited by 0SourcePDFScholar
2026

Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View Restoration

AAAI 2026technical

Multi-view diabetic retinopathy (DR) grading has achieved remarkable performance by capturing more comprehensive pathological features than single-view methods. However, complete multi-view fundus images are often difficult to obtain in clinical practice, and the performance degrades significantly w

Cited by 0SourcePDFScholar
2026

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

ICML 2026poster

Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's reality), watermarks trivially fail. Watermarks perturb output distributions away from the original, and in competitive mar…

Cited by 0SourceScholar
2026

MMPG: MoE-based Adaptive Multi-Perspective Graph Fusion for Protein Representation Learning

AAAI 2026technical

Graph Neural Networks (GNNs) have been widely adopted for Protein Representation Learning (PRL), as residue interaction networks can be naturally represented as graphs. Current GNN-based PRL methods typically rely on single-perspective graph construction strategies, which capture partial properties

Cited by 0SourcePDFScholar
2026

Prior Refinement Is Better: Diffusion-Driven Graph Harmonization for Federated Graph Learning

AAAI 2026technical

Federated Graph Learning (FGL) has emerged as a compelling paradigm for collaboratively training a global model while preserving the privacy of multi-source graphs. Nonetheless, FGL faces a critical challenge of data heterogeneity, where semantic and structural discrepancies across clients significa

Cited by 0SourcePDFScholar
2026

Sample Efficient Offline RL via T-Symmetry Enforced Latent State-Stitching

ICLR 2026poster

Offline reinforcement learning (RL) has achieved notable progress in recent years. However, most existing offline RL methods require a large amount of training data to achieve reasonable performance and offer limited out-of-distribution (OOD) generalization capability due to conservative data-relate…

Cited by 0SourceScholar
2026

Unifying Multi-View Knowledge for Graph Learning via Model Collaboration

AAAI 2026technical

With the increasing scale and complexity of graph data, node attributes are also becoming richer and more complex, particularly in the form of informative text. Classic GNNs equipped with shallow attribute encoders are no longer sufficient to handle such data independently, making model collaboratio

Cited by 0SourcePDFScholar
2025

Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus Diseases

AAAI 2025technical

With the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper…

Cited by 0SourcePDFScholar
2025

DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agent

ICLR 2025poster

On-device control agents, especially on mobile devices, are responsible for operating mobile devices to fulfill users' requests, enabling seamless and intuitive interactions. Integrating Multimodal Large Language Models (MLLMs) into these agents enhances their ability to understand and execute compl…

Cited by 13SourcePDFScholar
2025

Divide and Conquer: Coordinating Multiplex Mixture of Graph Learners to Handle Multi-Omics Analysis

IJCAI 2025

Graph learning has shown significant advantages in organizing and leveraging complex data, making it promising for numerous real-world applications with heterogeneous information, particularly multi-omics data analysis. Despite its potential in such scenarios, existing methods are still in their inf

Cited by 0SourcePDFScholar
2025

MYOPIA: Protecting Face Privacy from Malicious Personalized Text-to-Image Synthesis via Unlearnable Examples

AAAI 2025technical

Personalized text-to-image synthesis models, such as DreamBooth, have demonstrated significant potential in creating lifelike images tailored to a specific individual by fine-tuning from a limited set of face images and simple prompts. However, if misused, these model could pose a serious risk of pr…

2025

Mixture of Experts as Representation Learner for Deep Multi-View Clustering

AAAI 2025technical

Multi-view clustering (MVC) aims to integrate information from diverse data sources to facilitate the clustering process, which has achieved considerable success in various real-world applications. However, previous MVC methods typically employ one of two strategies: (1) designing separate feature e…

Cited by 0SourcePDFScholar
2025

Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors

IJCAI 2025

Integrating multi-omics datasets through data-driven analysis offers a comprehensive understanding of the complex biological processes underlying various diseases, particularly cancer. Graph Neural Networks (GNNs) have recently demonstrated remarkable ability to exploit relational structures in biol

Cited by 0SourcePDFScholar
2025

Refine then Classify: Robust Graph Neural Networks with Reliable Neighborhood Contrastive Refinement

AAAI 2025technical

Graph Neural Networks (GNNs) have exhibited remarkable capabilities for dealing with graph-structured data. However, recent studies have revealed their fragility to adversarial attacks, where imperceptible perturbations to the graph structure can easily mislead predictions. To enhance adversarial ro…

Cited by 0SourcePDFScholar
2025

SPA-BENCH: A COMPREHENSIVE BENCHMARK FOR SMARTPHONE AGENT EVALUATION

ICLR 2025spotlight

Smartphone agents are increasingly important for helping users control devices efficiently, with (Multimodal) Large Language Model (MLLM)-based approaches emerging as key contenders. Fairly comparing these agents is essential but challenging, requiring a varied task scope, the integration of agents…

2025

Strategy-Architecture Synergy: A Multi-View Graph Contrastive Paradigm for Consistent Representations

IJCAI 2025

Facing the growing diversity of multi-view data, multi-view graph-based models have made encouraging progress in handling multi-view data modeled as graphs. Graph Contrastive Learning (GCL) naturally fits multi-view graph data by treating their inherent views as augmentations. However, the developme

Cited by 0SourcePDFScholar
2025

Where Graph Meets Heterogeneity: Multi-View Collaborative Graph Experts

NeurIPS 2025poster

The convergence of graph learning and multi-view learning has propelled the emergence of multi-view graph neural networks (MGNNs), offering unprecedented capabilities to address complex real-world data characterized by heterogeneous yet interconnected information. While existing MGNNs exploit the p…

Cited by 0SourceScholar
2024

How to Learn Domain-Invariant Representations for Visual Reinforcement Learning: An Information-Theoretical Perspective

IJCAI 2024poster

Despite the impressive success in visual control challenges, Visual Reinforcement Learning (VRL) policies have struggled to generalize to other scenarios. Existing works attempt to empirically improve the generalization capability, lacking theoretical support. In this work, we explore how to learn d…

2024

OpticalDR: A Deep Optical Imaging Model for Privacy-Protective Depression Recognition

CVPR 2024poster

Depression Recognition (DR) poses a considerable challenge especially in the context of the growing concerns surrounding privacy. Traditional automatic diagnosis of DR technology necessitates the use of facial images undoubtedly expose the patient identity features and poses privacy risks. In order…

2024

What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return Prediction

AAAI 2024technical

In visual Reinforcement Learning (RL), the challenge of generalization to new environments is paramount. This study pioneers a theoretical analysis of visual RL generalization, establishing an upper bound on the generalization objective, encompassing policy divergence and Bellman error components. M…

2023

Beyond Graph Convolutional Network: An Interpretable Regularizer-Centered Optimization Framework

AAAI 2023technical

Graph convolutional networks (GCNs) have been attracting widespread attentions due to their encouraging performance and powerful generalizations. However, few work provide a general view to interpret various GCNs and guide GCNs' designs. In this paper, by revisiting the original GCN, we induce an in…

2023

DICNet: Deep Instance-Level Contrastive Network for Double Incomplete Multi-View Multi-Label Classification

AAAI 2023technical

In recent years, multi-view multi-label learning has aroused extensive research enthusiasm. However, multi-view multi-label data in the real world is commonly incomplete due to the uncertain factors of data collection and manual annotation, which means that not only multi-view features are often mis…

Cited by 56SourcePDFScholar
2023

Dual Low-Rank Graph Autoencoder for Semantic and Topological Networks

AAAI 2023technical

Due to the powerful capability to gather the information of neighborhood nodes, Graph Convolutional Network (GCN) has become a widely explored hotspot in recent years. As a well-established extension, Graph AutoEncoder (GAE) succeeds in mining underlying node representations via evaluating the quali…

Cited by 22SourcePDFScholar
2023

Graph Convolutional Kernel Machine versus Graph Convolutional Networks

NeurIPS 2023poster

Graph convolutional networks (GCN) with one or two hidden layers have been widely used in handling graph data that are prevalent in various disciplines. Many studies showed that the gain of making GCNs deeper is tiny or even negative. This implies that the complexity of graph data is often limited a…

2023

Highly Confident Local Structure Based Consensus Graph Learning for Incomplete Multi-View Clustering

CVPR 2023poster

Graph-based multi-view clustering has attracted extensive attention because of the powerful clustering-structure representation ability and noise robustness. Considering the reality of a large amount of incomplete data, in this paper, we propose a simple but effective method for incomplete multi-vie…

2023

Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RL

NeurIPS 2023poster

Offline reinforcement learning (RL) offers an appealing approach to real-world tasks by learning policies from pre-collected datasets without interacting with the environment. However, the performance of existing offline RL algorithms heavily depends on the scale and state-action space coverage of d…

2023

Masked Two-channel Decoupling Framework for Incomplete Multi-view Weak Multi-label Learning

NeurIPS 2023poster

Multi-view learning has become a popular research topic in recent years, but research on the cross-application of classic multi-label classification and multi-view learning is still in its early stages. In this paper, we focus on the complex yet highly realistic task of incomplete multi-view weak mu…

Cited by 18SourcePDFScholar
2022

Deep Object Detection with Example Attribute Based Prediction Modulation

ICASSP 2022accepted

Deep object detectors suffer from the gradient contribution imbalance during training. In this paper, we point out that such imbalance can be ascribed to the imbalance in example attributes, e.g., difficulty and shape variation degree. We further propose example attribute based prediction modulation…

Cited by 0SourceScholar