← Search

Zhen Cui

44 accepted papers

2026

Homophily-Heterogeneity Gradient Surgery for Federated Graph Learning

ICML 2026poster

Federated Graph Learning (FGL) facilitates privacy-preserving collaborative training of graph neural networks, yet homophily heterogeneity across subgraphs triggers optimization conflicts that degrade model generalization. Most existing solutions rely on multi-channel architectures to mitigate such …

Cited by 0SourceScholar
2026

Implicit Preference Alignment for Human Image Animation

ICML 2026poster

Human image animation has witnessed significant advancements, yet generating high-fidelity hand motions remains a persistent challenge due to their high degrees of freedom and motion complexity. While reinforcement learning from human feedback, particularly direct preference optimization, offers a p…

Cited by 0SourceScholar
2026

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

ICML 2026poster

Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploiting unlabeled image–text pairs. In this work, we propose Learning to Label, a rein…

Cited by 0SourceScholar
2026

Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

ICML 2026poster

Open-set supervised anomaly detection (OSAD) aims to identify unseen anomalies using limited anomalous supervision. However, existing prototype-based methods typically model normal data via a unimodal Gaussian prior, failing to capture inherent multi-modality and resulting in blurred decision bounda…

Cited by 0SourceScholar
2026

Parameter-Masked Decoupled Optimization for Cross-Domain Class-Incremental Learning

ICML 2026poster

Cross-domain class-incremental learning (CD-CIL) requires models to continuously acquire new classes across shifting domains while retaining previously learned knowledge. Existing approaches often entangle what to update with how to update, resulting in unstable adaptation and severe forgetting unde…

Cited by 0SourceScholar
2025

Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly Detection

CVPR 2025poster

In Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed anomaly samples, while overlooking critical priors of normal samples, leading to less effective discriminative boundaries. To address this issue,…

Cited by 0SourcePDFScholar
2025

Going Beyond Consistency: Target-oriented Multi-view Graph Neural Network

IJCAI 2025

Multi‐view learning has emerged as a pivotal research area driven by the growing heterogeneity of real‐world data, and graph neural network-based models, modeling multi-view data as multi-view graphs, have achieved remarkable performance by revealing its deep semantics. However, by assuming cross‐vi

2025

LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object Detection

ICCV 2025poster

Sparse annotation in remote sensing object detection poses significant challenges due to dense object distributions and category imbalances. Although existing Dense Pseudo-Label methods have demonstrated substantial potential in pseudo-labeling tasks, they remain constrained by selection ambiguities…

Cited by 0SourcePDFScholar
2025

Learn and Ensemble Bridge Adapters for Multi-domain Task Incremental Learning

NeurIPS 2025poster

Multi-domain task incremental learning (MTIL) demands models to master domain-specific expertise while preserving generalization capabilities. Inspired by human lifelong learning, which relies on revisiting, aligning, and integrating past experiences, we propose a Learning and Ensembling Bridge Ada…

Cited by 0SourceScholar
2025

M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation

ICASSP 2025accepted

The rapid growth of multimedia-sharing platforms drives the development of recommender systems. While traditional ID-based methods for mining user behavior signals are well-studied, research into multimodal sequential recommendation remains nascent. Current approaches face three critical challenges:…

Cited by 0SourceScholar
2025

Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised Detection

AAAI 2025technical

While existing semi-supervised object detection (SSOD) methods perform well in general scenes, they encounter challenges in handling oriented objects in aerial images. We experimentally find three gaps between general and oriented object detection in semi-supervised learning: 1) Sampling inconsist…

2025

One for All: Universal Topological Primitive Transfer for Graph Structure Learning

NeurIPS 2025poster

The non-Euclidean geometry inherent in graph structures fundamentally impedes cross-graph knowledge transfer. Drawing inspiration from texture transfer in computer vision, we pioneer topological primitives as transferable semantic units for graph structural knowledge. To address three critical barri…

Cited by 0SourceScholar
2025

Re-Attentional Controllable Video Diffusion Editing

AAAI 2025technical

Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resu…

2025

STDD: Spatio-Temporal Dual Diffusion for Video Generation

CVPR 2025poster

Diffusion probabilistic model is becoming the cornerstone of data generation, especially generating high-quality images. As an extension, video diffusion generation is in urgent need of a principled temporal-sequence diffusion way, while the spatial-domain diffusion dominates most video diffusion me…

Cited by 0SourcePDFScholar
2025

Scene Graph-Grounded Image Generation

AAAI 2025technical

With the beneft of explicit object-oriented reasoning capabilities of scene graphs, scene graph-to-image generation has made remarkable advancements in comprehending object coherence and interactive relations. Recent state-of-the-arts typically predict the scene layouts as an intermediate represent…

2025

UniHG: A Large-scale Universal Heterogeneous Graph Dataset and Benchmark for Representation Learning and Cross-Domain Transferring

NeurIPS 2025poster

Irregular data in the real world are usually organized as heterogeneous graphs consisting of multiple types of nodes and edges. However, current heterogeneous graph research confronts three fundamental challenges: i) Benchmark Deficiency, ii) Semantic Disalignment, and iii) Propagation Degradation.…

Cited by 0SourceScholar
2024

MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation

NeurIPS 2024poster

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general…

2024

Progressive Exploration-Conformal Learning for Sparsely Annotated Object Detection in Aerial Images

NeurIPS 2024poster

The ability to detect aerial objects with limited annotation is pivotal to the development of real-world aerial intelligence systems. In this work, we focus on a demanding but practical sparsely annotated object detection (SAOD) in aerial images, which encompasses a wider variety of aerial scenes wi…

Cited by 1SourcePDFScholar
2023

Deep Graph Structural Infomax

AAAI 2023technical

In the scene of self-supervised graph learning, Mutual Information (MI) was recently introduced for graph encoding to generate robust node embeddings. A successful representative is Deep Graph Infomax (DGI), which essentially operates on the space of node features but ignores topological structures,…

2023

Exploratory Inference Learning for Scribble Supervised Semantic Segmentation

AAAI 2023technical

Scribble supervised semantic segmentation has achieved great advances in pseudo label exploitation, yet suffers insufficient label exploration for the mass of unannotated regions. In this work, we propose a novel exploratory inference learning (EIL) framework, which facilitates efficient probing on…

Cited by 5SourcePDFScholar
2023

Progressive Bayesian Inference for Scribble-Supervised Semantic Segmentation

AAAI 2023technical

The scribble-supervised semantic segmentation is an important yet challenging task in the field of computer vision. To deal with the pixel-wise sparse annotation problem, we propose a Progressive Bayesian Inference (PBI) framework to boost the performance of the scribble-supervised semantic segmenta…

Cited by 3SourcePDFScholar
2023

Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly Detection

CVPR 2023poster

Weakly Supervised Video Anomaly Detection (WSVAD) is challenging because the binary anomaly label is only given on the video level, but the output requires snippet-level predictions. So, Multiple Instance Learning (MIL) is prevailing in WSVAD. However, MIL is notoriously known to suffer from many fa…

2022

CVNet: Contour Vibration Network for Building Extraction

CVPR 2022poster

The classic active contour model raises a great promising solution to polygon-based object extraction with the progress of deep learning recently. Inspired by the physical vibration theory, we propose a contour vibration network (CVNet) for automatic building boundary delineation. Different from the…

Cited by 22PDFcodeScholar
2021

Consistent Instance False Positive Improves Fairness in Face Recognition

CVPR 2021poster

Demographic bias is a significant challenge in practical face recognition systems. Several methods have been proposed to reduce the bias, which rely on accurate demographic annotations. However, such annotations are usually not available in real scenarios. Moreover, these methods are explicitly desi…

Cited by 66PDFcodeScholar
2021

Deep Wasserstein Graph Discriminant Learning for Graph Classification

AAAI 2021technical

Graph topological structures are crucial to distinguish different-class graphs. In this work, we propose a deep Wasserstein graph discriminant learning (WGDL) framework to learn discriminative embeddings of graphs in Wasserstein-metric (W-metric) matching space. In order to bypass the calculation of…

Cited by 19SourcePDFScholar
2021

Learning Normal Dynamics in Videos With Meta Prototype Network

CVPR 2021poster

Frame reconstruction (current or future frames) based on Auto-Encoder (AE) is a popular method for video anomaly detection. With models trained on the normal data, the reconstruction errors of anomalous scenes are usually much larger than those of normal ones. Previous methods introduced the memory…

Cited by 219PDFcodeScholar
2021

Scribble-Supervised Semantic Segmentation Inference

ICCV 2021poster

In this paper, we propose a progressive segmentation inference (PSI) framework to tackle with scribble-supervised semantic segmentation. In virtue of latent contextual dependency, we encapsulate two crucial cues, contextual pattern propagation and semantic label diffusion, to enhance and refine pixe…

Cited by 42PDFScholar
2021

Wasserstein Coupled Graph Learning for Cross-Modal Retrieval

ICCV 2021poster

Graphs play an important role in cross-modal image-text understanding as they characterize the intrinsic structure which is robust and crucial for the measurement of cross-modal similarity. In this work, we propose a Wasserstein Coupled Graph Learning (WCGL) method to deal with the cross-modal retri…

Cited by 29PDFScholar
2020

Cross-Modal Pattern-Propagation for RGB-T Tracking

CVPR 2020poster

Motivated by our observations on RGB-T data that pattern correlations are high-frequently recurred across modalities also along sequence frames, in this paper, we propose a cross-modal pattern-propagation (CMPP) tracking framework to diffuse instance patterns across RGB-T data on spatial domain as w…

Cited by 149PDFScholar
2020

Graph Wasserstein Correlation Analysis for Movie Retrieval

ECCV 2020poster

Movie graphs play an important role to bridge heterogenous modalities of videos and texts in human-centric retrieval. In this work, we propose Graph Wasserstein Correlation Analysis (GWCA) to deal with the core issue therein, i.e, cross heterogeneous graph comparison. Spectral graph filtering is int…

2020

Graph inference learning for semi-supervised classification

ICLR 2020poster

In this work, we address the semi-supervised classification of graph data, where the categories of those unlabeled nodes are inferred from labeled nodes as well as graph structures. Recent works often solve this problem with the advanced graph convolution in a conventional supervised manner, but the…

Cited by 39SourceScholar
2020

Pattern-Structure Diffusion for Multi-Task Learning

CVPR 2020poster

Inspired by the observation that pattern structures high-frequently recur within intra-task also across tasks, we propose a pattern-structure diffusion (PSD) framework to mine and propagate task-specific and task-across pattern structures in the task-level space for joint depth estimation, segmentat…

Cited by 111PDFScholar
2019

Pattern-Affinitive Propagation Across Depth, Surface Normal and Semantic Segmentation

CVPR 2019poster

In this paper, we propose a novel Pattern-Affinitive Propagation (PAP) framework to jointly predict depth, surface normal and semantic segmentation. The motivation behind it comes from the statistic observation that pattern-affinitive pairs recur much frequently across different tasks as well as wit…

Cited by 392PDFScholar
2018

Joint Task-Recursive Learning for Semantic Segmentation and Depth Estimation

ECCV 2018poster

In this paper, we propose a novel joint Task-Recursive Learning (TRL) framework for the closing-loop semantic segmentation and monocular depth estimation tasks. TRL can recursively refine the results of both tasks through serialized task-level interactions. In order to mutually-boost for each other,…

Cited by 261SourcePDFScholar