← Search

Jie Wen

61 accepted papers

2026

A Consensus Anchor-guided Hypergraph Framework For Incomplete Multi-view Clustering

ICML 2026poster

Handling large-scale incomplete multi-view data poses a significant challenge in unsupervised representation learning. While anchor-based strategies have alleviated computational burdens, they typically rely on shallow bipartite graphs restricted to pairwise relations, failing to capture complex hig…

Cited by 0SourceScholar
2026

Asymmetric Multi-View Clustering with Hyperbolic Uncertainty Modeling

ICML 2026spotlight

Deep Multi-View Clustering (MVC) aims to extract a unified semantic consensus from diverse data sources without supervision. However, current approaches relying on flat Euclidean embeddings often fail to model data uncertainty, resulting in rigid alignment where high-quality views are forced to drif…

Cited by 0SourceScholar
2026

Confident Block Diagonal Structure-Aware Invariable Graph Completion for Incomplete Multi-view Clustering

ICLR 2026poster

Multi-view clustering (MVC) adopts complementary information from multiple views to reveal the underlying structure of the data. However, the conventional MVC-based methods remain a crucial challenge on the incomplete multi-view clustering (IMVC) tasks, when some views of the multi-view data are mis…

Cited by 0SourceScholar
2026

Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification

CVPR 2026

While existing incomplete multi-view multi-label learning methods have achieved promising performance, few studies have focused on the issue of multi-view imbalance. Existing methods using gradient modulation or alternating optimization strategies alleviate this problem but often oversimplify the in

Cited by 0SourceScholar
2026

Detecting Fake News in Short Videos Through Multi-View Aggregation

AAAI 2026technical

The increasing prominence of short video platforms has positioned them as a primary channel for public awareness of current events, while also facilitating the widespread dissemination of fake news, thus highlighting the critical need for automated detection technologies. In contrast to fake news co

Cited by 0SourcePDFScholar
2026

Expandable, Compressible, Mineable: Open-World Thermal Infrared Image Restoration

ICML 2026poster

In open-world settings, thermal infrared (TIR) image degradations continuously emerge and evolve, while most existing all-in-one restoration methods are built on a closed-set assumption and struggle to continually adapt to novel degradations. To address this, we propose ECMRNet, an Expandable, Compr…

Cited by 0SourceScholar
2026

Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View Restoration

AAAI 2026technical

Multi-view diabetic retinopathy (DR) grading has achieved remarkable performance by capturing more comprehensive pathological features than single-view methods. However, complete multi-view fundus images are often difficult to obtain in clinical practice, and the performance degrades significantly w

Cited by 0SourcePDFScholar
2026

Permutation-Consistent Variational Encoding for Incomplete Multi-View Multi-Label Classification

ICLR 2026poster

Incomplete multi-view multi-label learning is fundamentally an information integration problem under simultaneous view and label incompleteness. We introduce Permutation-Consistent Variational Encoding framework (PCVE) with an information bottleneck strategy, which learns variational representations…

Cited by 0SourceScholar
2026

Prompt Tuning for CLIP on the Pretrained Manifold

ICML 2026poster

Prompt tuning introduces learnable prompt vectors that adapt pretrained vision-language models to downstream tasks in a parameter-efficient manner. However, under limited supervision, prompt tuning alters pretrained representations and drives downstream features away from the pretrained manifold tow…

Cited by 0SourceScholar
2026

Prototype-Based Semantic Consistency Alignment for Domain Adaptive Retrieval

AAAI 2026technical

Domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, enabling effective retrieval while mitigating domain discrepancies. However, existing methods encounter several fundamental limitations: 1) neglecting class-level semantic alignment and e

Cited by 0SourcePDFScholar
2026

Quality-aware and Soft Consistency Driven Representation Fusion for Incomplete Multi-view Multi-label Classification

AAAI 2026technical

Multi-view multi-label classification aims to utilize the rich information contained in multiple views for accurate classification. However, in real-world applications, its performance is often severely constrained by the concurrent missingness of both views and labels. To address this problem, this

Cited by 0SourcePDFScholar
2026

RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection

CVPR 2026

Underwater images commonly suffer from foreground-background ambiguity, loss of structural details, and severely reduced contrast, which collectively make underwater object detection (UOD) an inherently challenging task. To handle this issue, we present a residual-guided hierarchical calibration net

Cited by 0SourcecodeScholar
2026

S2C2Seg: Semantic-Spatial Consistency and Category Optimization for Open-Vocabulary Segmentation

CVPR 2026

Open-vocabulary semantic segmentation extends pixel-level recognition to arbitrary text-described categories. Despite strong global semantic understanding, vision-language models such as CLIP exhibit limited spatial precision and semantic ambiguity across large vocabularies, constraining their effec

Cited by 0SourceScholar
2026

SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything

CVPR 2026

Online 3D instance segmentation is a critical capability for embodied agents navigating in dynamic environments. However, a fundamental challenge remains in adapting powerful 2D foundation models, like SAM, to 3D online segmentation. Naively lifting SAM's 2D masks to 3D results in severe spatial fra

Cited by 0SourceScholar
2026

ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding

ICML 2026poster

Understanding and decoding brain activity into visual representations is a fundamental challenge at the intersection of neuroscience and artificial intelligence. While electroencephalogram (EEG) visual decoding has shown promise due to its non-invasive and low-cost nature, existing methods suffer fr…

Cited by 0SourceScholar
2025

A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers

NeurIPS 2025poster

Poison-only Clean-label Backdoor Attacks (PCBAs) aim to covertly inject attacker-desired behavior into DNNs by merely poisoning the dataset without changing the labels. To effectively implant a backdoor, multiple triggers are proposed for various attack requirements of Attack Success Rate (ASR) and…

Cited by 0SourceScholar
2025

ALRMR-GEC: Adjusting Learning Rate Based on Memory Rate to Optimize the Edit Scorer for Grammatical Error Correction

AAAI 2025technical

Edit-based approaches for Grammatical Error Correction (GEC) have attracted volume attention due to their outstanding explanations of the correction process and rapid inference. Through exploring the characteristics of the generalized and specific knowledge learning for GEC, we discover that efficie…

2025

AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs

ICASSP 2025accepted

Jailbreak vulnerabilities in Large Language Models (LLMs) refer to methods that extract malicious content from the model by carefully crafting prompts or suffixes, which has garnered significant attention from the research community. However, traditional attack methods, which primarily focus on the…

Cited by 0SourceScholar
2025

AdaptPFL: Unlocking Cross-Device Palmprint Recognition via Adaptive Personalized Federated Learning with Feature Decoupling

IJCAI 2025

Contactless palmprint recognition has recently emerged as a promising biometric technology. However, traditional methods that require sharing user data introduce substantial security risks. While federated learning offers privacy-preserving solutions, it often compromises recognition accuracy due to

Cited by 0SourcePDFScholar
2025

Confidence-Aware With Prototype Alignment for Partial Multi-label Learning

NeurIPS 2025poster

Label prototype learning has emerged as an effective paradigm in Partial Multi-Label Learning (PML), providing a distinctive framework for modeling structured representations of label semantics while naturally filtering noise through prototype-based label confidence estimation. However, existing pro…

Cited by 0SourceScholar
2025

Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus Diseases

AAAI 2025technical

With the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper…

Cited by 0SourcePDFScholar
2025

Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators

IJCAI 2025

Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this ch

2025

DiffusionREC: Diffusion Model with Adaptive Condition for Referring Expression Comprehension

AAAI 2025technical

The objective of referring expression comprehension (REC) is to accurately identify the object in an image described by a given expression. Existing REC methods, including transformer-based and graph-based approaches among others, have shown robust performance in REC tasks. In this study, we present…

Cited by 0SourcePDFScholar
2025

Efficient and Separate Authentication Image Steganography Network

ICML 2025spotlight

Image steganography hides multiple images for multiple recipients into a single cover image. All secret images are usually revealed without authentication, which reduces security among multiple recipients. It is elegant to design an authentication mechanism for isolated reception. We explore such me…

2025

Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-Training

IJCAI 2025

Multimodal protein features play a crucial role in protein function prediction. However, these features encompass a wide range of information, ranging from structural data and sequence features to protein attributes and interaction networks, making it challenging to decipher their complex interconne

2025

Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language Models

ICML 2025poster

With advancements in visual language models (VLMs) and large language models (LLMs), video anomaly detection (VAD) has progressed beyond binary classification to fine-grained categorization and multidimensional analysis. However, existing methods focus mainly on coarse-grained detection, lacking ano…

Cited by 0SourcePDFScholar
2025

Federated Incomplete Multi-view Clustering with Globally Fused Graph Guidance

ICML 2025poster

Federated multi-view clustering has been proposed to mine the valuable information within multi-view data distributed across different devices and has achieved impressive results while preserving the privacy. Despite great progress, most federated multi-view clustering methods only used global pseu…

2025

Federated Weakly Supervised Video Anomaly Detection with Multimodal Prompt

AAAI 2025technical

Video anomaly detection (VAD) aims at locating the abnormal events in videos. Recently, the Weakly Supervised VAD has made great progress, which only requires video-level annotations when training. In practical applications, different institutions may have different types of abnormal videos. However…

2025

FontAnimate: High Quality Few-shot Font Generation via Animating Font Transfer Process

ICCV 2025poster

Few-shot font generation (FFG) aims to create new font images by imitating the style from a limited set of reference images, while maintaining the content from the source images. Although this task has achieved significant progress, most existing methods still suffer from the incorrect generation of…

2025

Hierarchical Information Aggregation for Incomplete Multimodal Alzheimer's Disease Diagnosis

NeurIPS 2025poster

Alzheimer's Disease (AD) poses a significant health threat to the aging population, underscoring the critical need for early diagnosis to delay disease progression and improve patient quality of life. Recent advances in heterogeneous multimodal artificial intelligence (AI) have facilitated comprehen…

Cited by 0SourceScholar
2025

High-Confident Local Structure Guided Consensus Graph Learning For Incomplete Multi-view Clustering

IJCAI 2025

Current existing clustering methods for handling incomplete multi-view data primarily concentrate on learning a common representation or graph from the available views, while overlooking the latent information contained in the missing views and the imbalance of information among different views. Fur

2025

Incomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic Embedding

CVPR 2025poster

In incomplete multi-view multi-label learning scenarios, it is crucial to use the incomplete multi-view data to extract consistent and specific representations from different data sources and to fully exploit the missing label information. However, most previous approaches ignore the separation prob…

Cited by 0SourcePDFScholar
2025

Learning Compact Semantic Information for Incomplete Multi-View Missing Multi-Label Classification

ICML 2025poster

Multi-view data involves various data forms, such as multi-feature, multi-sequence and multimodal data, providing rich semantic information for downstream tasks. The inherent challenge of incomplete multi-view missing multi-label learning lies in how to effectively utilize limited supervision and in…

Cited by 0SourcePDFScholar
2025

Learning from Disjoint Views: A Contrastive Prototype Matching Network for Fully Incomplete Multi-View Clustering

NeurIPS 2025poster

Multi-view clustering aims to enhance clustering performance by leveraging information from diverse sources. However, its practical application is often hindered by a barrier: the lack of correspondences across views. This paper focuses on the understudied problem of fully incomplete multi-view clus…

Cited by 0SourceScholar
2025

Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval

AAAI 2025technical

Deep online cross-modal hashing has gained much attention from researchers recently, as its promising applications with low storage requirement, fast retrieval efficiency and cross modality adaptive, etc. However, there still exists some technical hurdles that hinder its applications, e.g., 1) how t…

Cited by 0SourcePDFScholar
2025

Multi-view Evidential Learning-based Medical Image Segmentation

AAAI 2025technical

Medical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limi…

Cited by 0SourcePDFScholar
2025

NeuroH-TGL: Neuro-Heterogeneity Guided Temporal Graph Learning Strategy for Brain Disease Diagnosis

NeurIPS 2025poster

Dynamic functional brain networks (DFBNs) are powerful tools in neuroscience research. Recent studies reveal that DFBNs contain heterogeneous neural nodes with more extensive connections and more drastic temporal changes, which play pivotal roles in coordinating the reorganization of the brain. More…

Cited by 0SourceScholar
2025

Omni-Dimensional State Space Model-driven SAM for Pixel-level Anomaly Detection

IJCAI 2025

Pixel-level anomaly detection is indispensable in industrial defect detection and medical diagnosis. Recently, Segment Anything Model (SAM) has achieved promising results in many vision tasks. However, direct application of the SAM to pixel-level anomaly detection tasks results in unsatisfactory per

Cited by 0SourcePDFScholar
2025

Towards VLM-based Hybrid Explainable Prompt Enhancement for Zero-Shot Industrial Anomaly Detection

IJCAI 2025

Zero-Shot Industrial Anomaly Detection (ZSIAD) aims to identify and localize anomalies in industrial images from unseen categories. Owing to the powerful generalization capabilities, Vision-Language Models (VLMs) have achieved growing interest in ZSIAD. To guide the model toward understanding and lo

Cited by 0SourcePDFScholar
2025

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

NeurIPS 2025poster

Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this pa…

Cited by 0SourcecodeScholar
2025

Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency Learning

ICCV 2025poster

To reduce the reliance of visible-infrared person re-identification (ReID) models on labeled cross-modal samples, this paper explores a weakly supervised cross-modal person ReID method that uses only single-modal sample identity labels, addressing scenarios where cross-modal identity labels are unav…

2024

A Two-Stage Information Extraction Network for Incomplete Multi-View Multi-Label Classification

AAAI 2024technical

Recently, multi-view multi-label classification (MvMLC) has received a significant amount of research interest and many methods have been proposed based on the assumptions of view completion and label completion. However, in real-world scenarios, multi-view multi-label data tends to be incomplete du…

2024

Attention-Induced Embedding Imputation for Incomplete Multi-View Partial Multi-Label Classification

AAAI 2024technical

As a combination of emerging multi-view learning methods and traditional multi-label classification tasks, multi-view multi-label classification has shown broad application prospects. The diverse semantic information contained in heterogeneous data effectively enables the further development of mult…

Cited by 13SourcePDFScholar
2024

Batch Singular Value Polarization and Weighted Semantic Augmentation for Universal Domain Adaptation

ICML 2024poster

As a more challenging domain adaptation setting, universal domain adaptation (UniDA) introduces category shift on top of domain shift, which needs to identify unknown category in the target domain and avoid misclassifying target samples into source private categories. To this end, we propose a novel…

Cited by 0SourcePDFScholar
2024

Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering Structures

AAAI 2024technical

Incomplete multi-view clustering (IMVC) aims to reveal shared clustering structures within multi-view data, where only partial views of the samples are available. Existing IMVC methods primarily suffer from two issues: 1) Imputation-based methods inevitably introduce inaccurate imputations, which in…

Cited by 16SourcePDFScholar
2024

Diffusion-based Missing-view Generation With the Application on Incomplete Multi-view Clustering

ICML 2024poster

As a branch of clustering, multi-view clustering has received much attention in recent years. In practical applications, a common phenomenon is that partial views of some samples may be missing in the collected multi-view data, which poses a severe challenge to design the multi-view learning model a…

Cited by 3SourcePDFScholar
2024

Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion Models

CVPR 2024poster

Few-shot font generation (FFG) produces stylized font images with a limited number of reference samples which can significantly reduce labor costs in manual font designs. Most existing FFG methods follow the style-content disentanglement paradigm and employ the Generative Adversarial Network (GAN) t…

2024

HACDR-Net: Heterogeneous-Aware Convolutional Network for Diabetic Retinopathy Multi-Lesion Segmentation

AAAI 2024technical

Diabetic Retinopathy (DR), the leading cause of blindness in diabetic patients, is diagnosed by the condition of retinal multiple lesions. As a difficult task in medical image segmentation, DR multi-lesion segmentation faces the main concerns as follows. On the one hand, retinal lesions vary in loca…

2024

Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image Recognition

ICML 2024poster

Large-scale pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities in image recognition tasks. Recent approaches typically employ supervised fine-tuning methods to adapt CLIP for zero-shot multi-label image recognition tasks. However, obtaining sufficient…

Cited by 3SourcePDFScholar
2024

Long Short-Term Dynamic Prototype Alignment Learning for Video Anomaly Detection

IJCAI 2024poster

Video anomaly detection (VAD) is the core problem of intelligent video surveillance. Previous methods commonly adopt the unsupervised paradigm of frame reconstruction or prediction. However, the lack of mining of temporal dependent relationships and diversified event patterns within videos limit the…

Cited by 6SourcePDFScholar
2024

Optimal Transport-based Labor-free Text Prompt Modeling for Sketch Re-identification

NeurIPS 2024poster

Sketch Re-identification (Sketch Re-ID), which aims to retrieve target person from an image gallery based on a sketch query, is crucial for criminal investigation, law enforcement, and missing person searches. Existing methods aim to alleviate the modality gap by employing semantic metrics constra…

Cited by 0SourcePDFScholar
2024

Partial Multi-View Multi-Label Classification via Semantic Invariance Learning and Prototype Modeling

ICML 2024poster

The difficulty of partial multi-view multi-label learning lies in coupling the consensus of multi-view data with the task relevance of multi-label classification, under the condition where partial views and labels are unavailable. In this paper, we seek to compress cross-view representation to maxim…

Cited by 2SourcePDFScholar
2023

DICNet: Deep Instance-Level Contrastive Network for Double Incomplete Multi-View Multi-Label Classification

AAAI 2023technical

In recent years, multi-view multi-label learning has aroused extensive research enthusiasm. However, multi-view multi-label data in the real world is commonly incomplete due to the uncertain factors of data collection and manual annotation, which means that not only multi-view features are often mis…

Cited by 56SourcePDFScholar
2023

Highly Confident Local Structure Based Consensus Graph Learning for Incomplete Multi-View Clustering

CVPR 2023poster

Graph-based multi-view clustering has attracted extensive attention because of the powerful clustering-structure representation ability and noise robustness. Considering the reality of a large amount of incomplete data, in this paper, we propose a simple but effective method for incomplete multi-vie…

2023

Incomplete Multi-View Multi-Label Learning via Label-Guided Masked View- and Category-Aware Transformers

AAAI 2023technical

As we all know, multi-view data is more expressive than single-view data and multi-label annotation enjoys richer supervision information than single-label, which makes multi-view multi-label learning widely applicable for various pattern recognition tasks. In this complex representation learning pr…

2023

Masked Two-channel Decoupling Framework for Incomplete Multi-view Weak Multi-label Learning

NeurIPS 2023poster

Multi-view learning has become a popular research topic in recent years, but research on the cross-application of classic multi-label classification and multi-view learning is still in its early stages. In this paper, we focus on the complex yet highly realistic task of incomplete multi-view weak mu…

Cited by 18SourcePDFScholar
2023

Tensorized Incomplete Multi-View Clustering with Intrinsic Graph Completion

AAAI 2023technical

Most of the existing incomplete multi-view clustering (IMVC) methods focus on attaining a consensus representation from different views but ignore the important information hidden in the missing views and the latent intrinsic structures in each view. To tackle these issues, in this paper, a unified…

2022

Deep Object Detection with Example Attribute Based Prediction Modulation

ICASSP 2022accepted

Deep object detectors suffer from the gradient contribution imbalance during training. In this paper, we point out that such imbalance can be ascribed to the imbalance in example attributes, e.g., difficulty and shape variation degree. We further propose example attribute based prediction modulation…

Cited by 0SourceScholar
2021

Scalable Discriminative Discrete Hashing For Large-Scale Cross-Modal Retrieval

ICASSP 2021accepted

Cross-modal hashing has received increasing research attentions due to its less storage and efficient retrieval. However, most existing cross-modal hashing methods focus only on exploring multi-modal information, while underestimate the significance of local and Euclidean structure information on th…

Cited by 0SourceScholar
2021

Unified Tensor Framework for Incomplete Multi-view Clustering and Missing-view Inferring

AAAI 2021technical

In this paper, we propose a novel method, referred to as incomplete multi-view tensor spectral clustering with missing-view inferring (IMVTSC-MVI) to address the challenging multi-view clustering problem with missing views. Different from the existing methods which commonly focus on exploring the ce…

Cited by 157SourcePDFScholar
2020

CDIMC-net: Cognitive Deep Incomplete Multi-view Clustering Network

IJCAI 2020poster

In recent years, incomplete multi-view clustering, which studies the challenging multi-view clustering problem on missing views, has received growing research interests. Although a series of methods have been proposed to address this issue, the following problems still exist: 1) Almost all of the ex…

Cited by 0SourcePDFScholar