← Search

Mang Ye

104 accepted papers

2026

API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment Analysis

ICML 2026poster

Multimodal sentiment analysis aims to infer human emotions by integrating signals from diverse modalities. However, missing modalities are common in real-world applications due to sensor failure, data corruption, or privacy concerns. Existing approaches typically follow two main paradigms: recovery-…

Cited by 0SourceScholar
2026

Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attack

CVPR 2026

Federated Learning (FL), a distributed learning paradigm that enables local training on user-held data across decentralized devices, is vulnerable to backdoor attacks due to limited visibility into client updates. Exploiting this opacity, adversaries induce targeted misbehavior on trigger inputs wit

Cited by 0SourcecodeScholar
2026

Cross-Modal Semantic Decoupling and Transfer for Text-to-Visible-Infrared Person Re-Identification

ICML 2026poster

Text-to-Image Person Re-Identification (TI-ReID) retrieves visible pedestrian images using text queries. Yet in low-light or nighttime settings, visible images lack sufficient identity details, while infrared images effectively capture pedestrian contours and textures. To enable all-day surveillance…

Cited by 0SourceScholar
2026

Divide, Conquer and Unite: Hierarchical Style-Recalibrated Prototype Alignment for Federated Medical Segmentation

AAAI 2026technical

Federated learning enables multiple medical institutions to train a global model without sharing data, yet feature heterogeneity from diverse scanners or protocols remains a major challenge. Many existing works attempt to address this issue by leveraging model representations (e.g., mean feature vec

Cited by 0SourcePDFScholar
2026

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

CVPR 2026

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches based on supervised fine-tuning often suffer from limited generalization and poor i

Cited by 0SourcecodeScholar
2026

Enhancing Cross-subject Emotion Recognition via Heterogeneous Distribution Augmentation and Collaborative Learning

ICML 2026poster

Cross-subject emotion recognition aims to improve a model's generalization to previously unseen subjects. Existing methods are mainly built upon domain generalization or data augmentation, but suffer from two major limitations: 1) heavy dependence on modality-specific feature designs—almost exclusiv…

Cited by 0SourceScholar
2026

FedPissa: Towards Federated Personalized Adaptation of Foundation Models via LoRA Subspace Mapping

ICML 2026spotlight

LoRA efficiently adapts large pre-trained models via low-rank updates, making it a strong parameter-efficient fine-tuning (PEFT) method. When integrated with Federated Learning (FL), it enables collaborative fine-tuning across distributed clients, leveraging rich downstream data without exposing pri…

Cited by 0SourceScholar
2026

FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction

CVPR 2026

Federated Graph Learning (FGL) has emerged as a principled framework for decentralized training of Graph Neural Networks (GNNs) while preserving data privacy. In subgraph-FL scenarios, however, structural noise arising from data collection and storage can damage the GNN message-passing scheme of cli

Cited by 0SourcecodeScholar
2026

Interactive Person Retrieval via Multi-Turn Multimodal Conversation

ICML 2026poster

Traditional text-based person retrieval approaches typically rely on single-shot textual queries, which are generally incomplete or vague in real-world scenarios. Recently, chat-based person retrieval methods enable iterative query refinement via question-answering interactions between the system an…

Cited by 0SourceScholar
2026

Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement

ICLR 2026poster

Multimodal editing large models have demonstrated powerful editing capabilities across diverse tasks. However, a persistent and long-standing limitation is the decline in facial identity (ID) consistency during realistic portrait editing. Due to the human eye’s high sensitivity to facial features, s…

Cited by 0SourcecodeScholar
2026

Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language Model

AAAI 2026technical

The massive scale of data and computation required for training Multimodal Large Language Models (MLLMs) has fueled the rise of Fine-Tuning as a Service (FTaaS), enabling users to rapidly customize models for diverse real-world tasks. While FTaaS democratizes access to advanced multimodal intelligen

Cited by 0SourcePDFScholar
2026

RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation

AAAI 2026technical

Humanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate extensive training datasets to achieve robust multimodal decision-making capabilities and generalizable visuomotor control.

Cited by 0SourcePDFScholar
2026

Rethinking Federated Prompt Learning for Medical Images: From Textual Tuning to Visual Manifold Anchoring

ICML 2026poster

Federated Prompt Learning (FPL) adapts Vision-Language Models to privacy-sensitive medical imaging, typically via a textual tuning paradigm that assumes the frozen visual encoder provides a discriminative feature geometry. We argue this assumption breaks down in medical settings, leading to two geom…

Cited by 0SourceScholar
2026

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

CVPR 2026

Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new compositional safety risks that emerge from complex text-image interactions.Such cross-modal couplings can produce unsafe semantics

Cited by 0SourcecodeScholar
2026

Shift-Dependent Asymmetry: Orthogonal Inverse Low-Rank Adaptation for Federated Medical Segmentation

ICML 2026poster

Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of segmentation foundation models for medical imaging. However, most federated LoRA methods adopt a uniform aggregation rule, which breaks under the encoder–decoder asymmetry in medical segmentation: the encoder is dominated by appea…

Cited by 0SourceScholar
2026

The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution

ICML 2026poster

Large Reasoning Models (LRMs) enhance performance by generating explicit Chain-of-Thought (CoT) trajectories, yet enabling them to self-evaluate correctness without external supervision remains a critical challenge. Existing methods often rely on ground-truth labels or shallow output probabilities, …

Cited by 0SourceScholar
2026

Towards Cross-Modal Preservation, Consistency and Alignment for Privacy-Preserving Visible-Infrared Person Re-Identification

CVPR 2026

Privacy-preserving Person Re-Identification (PP-ReID) addresses the core privacy-utility trade-off in Re-ID by retrieving a person across multiple non-overlapping cameras while applying anonymization techniques to protect sensitive information. However, prior PP-ReID studies are confined to single-m

Cited by 0SourcecodeScholar
2026

Towards Realistic Lifelong Re-identification: Identity Recurrence with Changing Clothes

ICML 2026poster

Existing lifelong person re-identification (Re-ID) methods assume that each identity maintains a relatively stable appearance distribution over time. However, in real-world scenarios, identities often reappear asynchronously with substantial clothing changes, which is not modeled in existing lifelon…

Cited by 0SourceScholar
2026

Towards Robust Text-Attributed Federated Graph Learning: Multimodal Threats and Defense

AAAI 2026technical

Text-Attributed Graphs (TAGs) are graphs where both nodes and edges are associated with text attributes. To leverage their semantic richness, recent efforts have integrated large language models (LLMs) with graph neural networks, leading to the development of GraphLLMs. However, many real-world data

Cited by 0SourcePDFScholar
2026

WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification

CVPR 2026

Recent person re-identification (ReID) leverages heterogeneous sensing with multiple modalities and viewpoints to improve robustness across diverse conditions. However, most approaches target predefined scenario pairs (e.g., visible-infrared or aerial-ground) and train separate task-specific models.

Cited by 0SourcecodeScholar
2025

$S^2$FGL: Spatial Spectral Federated Graph Learning

ICML 2025poster

Federated Graph Learning (FGL) combines the privacy-preserving capabilities of federated learning (FL) with the strong graph modeling capability of Graph Neural Networks (GNNs). Current research addresses subgraph-FL only from the structural perspective, neglecting the propagation of graph signals o…

2025

An Empirical Study of Federated Prompt Learning for Vision Language Model

IJCAI 2025

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. Th

2025

Backdoor Cleaning without External Guidance in MLLM Fine-tuning

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious security risks, as malicious fine-tuning can implant backdoor…

Cited by 0SourcecodeScholar
2025

Be Confident: Uncovering Overfitting in MLLM Multi-Task Tuning

ICML 2025poster

Fine-tuning Multimodal Large Language Models (MLLMs) in multi-task learning scenarios has emerged as an effective strategy for achieving cross-domain specialization. However, multi-task fine-tuning frequently induces performance degradation on open-response datasets. We posit that free-form answer g…

Cited by 0SourcePDFScholar
2025

CAN: Leveraging Clients As Navigators for Generative Replay in Federated Continual Learning

ICML 2025poster

Generative replay (GR) has been extensively validated in continual learning as a mechanism to synthesize data and replay past knowledge to mitigate forgetting. By leveraging synthetic rather than real data for the replay, GR has been adopted in some federated continual learning (FCL) approaches to…

Cited by 0SourcePDFScholar
2025

Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models

ICML 2025spotlight

Multimodal large language models (MLLMs) have achieved impressive progress in tasks such as visual question answering and visual understanding, but they still face significant challenges in emotional reasoning. Current methods to enhance emotional understanding typically rely on fine-tuning or manua…

Cited by 0SourcePDFScholar
2025

Chat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignment

CVPR 2025poster

Traditional text-based person retrieval (TPR) relies on a single-shot text as query to retrieve the target person, assuming that the query completely captures the user's search intent. However, in real-world scenarios, it can be challenging to ensure the information completeness of such single-shot…

2025

Cheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identification

CVPR 2025poster

Person re-identification (ReID) is the task of matching individuals across different camera views. Existing approaches typically employ neural networks to extract discriminative features, ranking gallery images based on their similarities to probe images. While effective, these methods are often enh…

2025

DKDR: Dynamic Knowledge Distillation for Reliability in Federated Learning

NeurIPS 2025poster

Federated Learning (FL) has demonstrated a promising future in privacy-friendly collaboration but it faces the data heterogeneity problem. Knowledge Distillation (KD) can serve as an effective method to address this issue. However, challenges arise from the unreliability of existing distillation met…

Cited by 0SourcecodeScholar
2025

EAGLES: Towards Effective, Efficient, and Economical Federated Graph Learning via Unified Sparsification

ICML 2025poster

Federated Graph Learning (FGL) has gained significant attention as a privacy-preserving approach to collaborative learning, but the computational demands increase substantially as datasets grow and Graph Neural Network (GNN) layers deepen. To address these challenges, we propose $\textbf{EAGLES}$, a…

Cited by 0SourcePDFScholar
2025

EMOE: Modality-Specific Enhanced Dynamic Emotion Experts

CVPR 2025poster

Multimodal Emotion Recognition (MER) aims to predict human emotions by leveraging multiple modalities, such as vision, acoustics, and language. However, due to the heterogeneity of these modalities, MER faces two key challenges: modality balance dilemma and modality specialization disappearance. Exi…

2025

Energy-based Backdoor Defense Against Federated Graph Learning

ICLR 2025oral

Federated Graph Learning is rapidly evolving as a privacy-preserving collaborative approach. However, backdoor attacks are increasingly undermining federated systems by injecting carefully designed triggers that lead to the model making incorrect predictions. Trigger structures and injection locatio…

Cited by 0SourcePDFScholar
2025

FedPHA: Federated Prompt Learning for Heterogeneous Client Adaptation

ICML 2025poster

Federated Prompt Learning (FPL) adapts pre-trained Vision-Language Models (VLMs) to federated learning through prompt tuning, leveraging their transferable representations and strong generalization capabilities. Traditional methods often require uniform prompt lengths for federated aggregation, limi…

Cited by 0SourcePDFScholar
2025

FedSPA: Generalizable Federated Graph Learning under Homophily Heterogeneity

CVPR 2025poster

Federated Graph Learning (FGL) has emerged as a solution to address real-world privacy concerns and data silos in graph learning, which relies on Graph Neural Networks (GNNs). Nevertheless, the homophily level discrepancies within the local graph data of clients, termed homophily heterogeneity, sign…

2025

Federated Disentangled Tuning with Textual Prior Decoupling and Visual Dynamic Adaptation

ICML 2025poster

Federated Parameter-Efficient Fine-Tuning aims to adapt Vision-Language Models for downstream tasks in distributed environments. However, data heterogeneity across participants hinders collaborative effectiveness, necessitating personalized adaptation to cover distinct data distributions. Current pe…

2025

GHOST: Generalizable One-Shot Federated Graph Learning with Proxy-Based Topology Knowledge Retention

ICML 2025poster

Federated Graph Learning (FGL) proposes an effective approach to collaboratively training Graph Neural Networks (GNNs) while maintaining privacy. Nevertheless, communication efficiency becomes a critical bottleneck in environments with limited resources. In this context, one-shot FGL emerges as a pr…

2025

Image-assisted Label Connective Completion for Vessel Segmentation with Insufficient Annotations

ICASSP 2025accepted

Automatic and accurate vessel segmentation is crucial for disease diagnosis. Deep learning methods are widely used, but their promising results rely on accurately annotated data. Due to complex vessel morphology and low-contrast image, accurate vessel delineation poses a practical challenge, resulti…

Cited by 0SourceScholar
2025

Label-Free Backdoor Attacks in Vertical Federated Learning

AAAI 2025technical

Vertical Federated Learning (VFL) involves multiple clients collaborating to train a global model, with distributed features of shared samples. While it becomes a critical privacy-preserving learning paradigm, its security can be significantly compromised by backdoor attacks, where a malicious clien…

2025

Learn from Downstream and Be Yourself in Multimodal Large Language Models Fine-Tuning

ICML 2025poster

Multimodal Large Language Model (MLLM) has demonstrated strong generalization capabilities across diverse distributions and tasks, largely due to extensive pre-training datasets. Fine-tuning MLLM has become a common practice to improve performance on specific downstream tasks. However, during fine-t…

Cited by 9SourcePDFScholar
2025

LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models

CVPR 2025poster

While Multimodal Large Language Models (MLLMs) excel at generalizing across modalities and tasks, effectively adapting them to specific downstream tasks while simultaneously retaining both general and specialized knowledge remains challenging. Although Low-Rank Adaptation (LoRA) is widely used to ef…

2025

MARS-VFL: A Unified Benchmark for Vertical Federated Learning with Realistic Evaluation

NeurIPS 2025spotlight

Vertical Federated Learning (VFL) has emerged as a critical privacy-preserving learning paradigm, enabling collaborative model training by leveraging distributed features across clients. However, due to privacy concerns, there are few publicly available real-world datasets for evaluating VFL methods…

Cited by 0SourceScholar
2025

MDFG: Multi-Dimensional Fine-Grained Modeling for Fatigue Detection

AAAI 2025technical

Fatigue is a critical factor contributing to accidents in industries such as safety monitoring and engineering construction. Fatigue exhibits dynamic complexity and non-stationary characteristics, so there are many intermediate states of short-term variation between alert and fatigue. Capturing and…

2025

MOTION: Multi-Sculpt Evolutionary Coarsening for Federated Continual Graph Learning

NeurIPS 2025poster

Graph neural networks (GNNs) have achieved remarkable success in various domains but typically rely on centralized, static graphs, which limits their applicability in distributed, evolving environments. To address this limitation, we define the task of Federated Continual Graph Learning (FCGL), a pa…

Cited by 0SourceScholar
2025

MoodAngels: A Retrieval-augmented Multi-agent Framework for Psychiatry Diagnosis

NeurIPS 2025poster

The application of AI in psychiatric diagnosis faces significant challenges, including the subjective nature of mental health assessments, symptom overlap across disorders, and privacy constraints limiting data availability. To address these issues, we present MoodAngels, the first specialized multi…

Cited by 0SourceScholar
2025

Multi-order Orchestrated Curriculum Distillation for Model-Heterogeneous Federated Graph Learning

NeurIPS 2025poster

Federated Graph Learning (FGL) has been shown to be particularly effective in enabling collaborative training of Graph Neural Networks (GNNs) in decentralized settings. Model-heterogeneous FGL further enhances practical applicability by accommodating client preferences for diverse model architecture…

Cited by 0SourceScholar
2025

NightReID: A Large-Scale Nighttime Person Re-Identification Benchmark

AAAI 2025technical

Person re-identification (Re-ID) is crucial for intelligent surveillance systems, facilitating the identification of individuals across multiple camera views. While significant advancements have been made for daytime scenarios, ensuring reliable Re-ID performance during nighttime remains a significa…

2025

OASIS: One-Shot Federated Graph Learning via Wasserstein Assisted Knowledge Integration

NeurIPS 2025poster

Federated Graph Learning (FGL) offers a promising framework for collaboratively training Graph Neural Networks (GNNs) while preserving data privacy. In resource-constrained environments, One-shot Federated Learning (OFL) emerges as an effective solution by limiting communication to a single round. C…

Cited by 0SourceScholar
2025

PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric Network

NeurIPS 2025poster

With the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules…

Cited by 0SourcecodeScholar
2025

Pixel-wise Divide and Conquer for Federated Vessel Segmentation

IJCAI 2025

Accurate vessel segmentation is essential for diagnosing and managing vascular and ophthalmic diseases. Traditional learning-based vessel segmentation methods heavily rely on high-quality, pixel-level annotated datasets. However, segmentation performance suffers significantly when applied in federat

Cited by 0SourcePDFScholar
2025

Prototype-guided Knowledge Propagation with Adaptive Learning for Lifelong Person Re-identification

IJCAI 2025

Lifelong Person Re-identification (LReID) is essential in dynamic camera networks, which continually adapts to new environments while preserving previously acquired knowledge. Existing LReID techniques often preserve samples from past datasets to maintain old knowledge, potentially leading to privac

2025

Rethinking Fair Federated Learning from Parameter and Client View

NeurIPS 2025poster

Federated Learning is a promising technique that enables collaborative machine learning while preserving participant privacy. With respect to multi-party collaboration, achieving performance fairness acts as a critical challenge in federated systems. Existing explorations mainly focus on considering…

Cited by 0SourcecodeScholar
2025

SPMC: Self-Purifying Federated Backdoor Defense via Margin Contribution

ICML 2025poster

Federated Learning (FL) enables collaborative training with privacy preservation but is vulnerable to backdoor attacks, where malicious clients degrade model performance on targeted inputs. These attacks exploit FL decentralized nature, while existing defenses, based on isolated behaviors and fixed…

2025

Splitting with Importance-aware Updating for Heterogeneous Federated Learning with Large Language Models

ICML 2025poster

Federated learning provides an efficient privacy-preserving distributed training framework for large language models, addressing the growing scarcity of publicly available training data while enabling the utilization of private datasets. While integrating large language model fine-tuning with federa…

2025

Synthetic Data is an Elegant GIFT for Continual Vision-Language Models

CVPR 2025poster

Pre-trained Vision-Language Models (VLMs) require Continual Learning (CL) to efficiently update their knowledge and adapt to various downstream tasks without retraining from scratch. However, for VLMs, in addition to the loss of knowledge previously learned from downstream tasks, pre-training knowle…

2025

TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-Identification

AAAI 2025technical

Unsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo la…

2025

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

IJCAI 2025

In real applications, person re-identification (ReID) expects to retrieve the target person at any time, including both daytime and nighttime, ranging from short-term to long-term. However, existing ReID tasks and datasets cannot meet this requirement, as they are constrained by available time and o

2025

Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-Identification

NeurIPS 2025spotlight

In object re-identification (ReID) task, both cross-modal and multi-modal retrieval methods have achieved notable progress. However, existing approaches are designed for specific modality and category (person or vehicle) retrieval task, lacking generalizability to others. Acquiring multiple task-spe…

Cited by 0SourcecodeScholar
2025

Uncertain Multimodal Intention and Emotion Understanding in the Wild

CVPR 2025poster

Understanding intention and emotion from social media poses unique challenges due to the inherent uncertainty in multimodal data, where posts often contain incomplete or missing modalities. While this uncertainty reflects real-world scenarios, it remains underexplored within the computer vision comm…

2025

Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

ICCV 2025poster

Unsupervised visible-infrared person re-identification (USL-VI-ReID) aims to train a cross-modality retrieval model without labels, reducing the reliance on expensive cross-modality manual annotation. However, existing USL-VI-ReID methods rely on artificially cross-modality paired data as implicit s…

2024

An Empirical Study of CLIP for Text-Based Person Search

AAAI 2024technical

Text-based Person Search (TBPS) aims to retrieve the person images using natural language descriptions. Recently, Contrastive Language Image Pretraining (CLIP), a universal large cross-modal vision-language pre-training model, has remarkably performed over various cross-modal downstream tasks due to…

2024

Contextual Augmented Global Contrast for Multimodal Intent Recognition

CVPR 2024poster

Multimodal intent recognition (MIR) aims to perceive the human intent polarity via language visual and acoustic modalities. The inherent intent ambiguity makes it challenging to recognize in multimodal scenarios. Existing MIR methods tend to model the individual video independently ignoring global c…

Cited by 11SourcePDFScholar
2024

DifTraj: Diffusion Inspired by Intrinsic Intention and Extrinsic Interaction for Multi-Modal Trajectory Prediction

IJCAI 2024poster

Recent years have witnessed the success of generative adversarial networks and diffusion models in multi-model trajectory prediction. However, prevailing algorithms only explicitly consider human interaction, but ignore the modeling of human intention, yielding that the generated results deviate lar…

Cited by 1SourcePDFScholar
2024

Empowering Visible-Infrared Person Re-Identification with Large Foundation Models

NeurIPS 2024poster

Visible-Infrared Person Re-identification (VI-ReID) is a challenging cross-modal retrieval task due to significant modality differences, primarily resulting from the absence of color information in the infrared modality. The development of large foundation models like Large Language Models (LLMs) an…

Cited by 3SourcePDFScholar
2024

Fair Federated Learning under Domain Skew with Local Consistency and Domain Diversity

CVPR 2024poster

Federated learning (FL) has emerged as a new paradigm for privacy-preserving collaborative training. Under domain skew the current FL approaches are biased and face two fairness problems. 1) Parameter Update Conflict: data disparity among clients leads to varying parameter importance and inconsisten…

Cited by 21SourcePDFScholar
2024

FedSSP: Federated Graph Learning with Spectral Knowledge and Personalized Preference

NeurIPS 2024poster

Personalized Federated Graph Learning (pFGL) facilitates the decentralized training of Graph Neural Networks (GNNs) without compromising privacy while accommodating personalized requirements for non-IID participants. In cross-domain scenarios, structural heterogeneity poses significant challenges fo…

2024

Federated Graph Learning under Domain Shift with Generalizable Prototypes

AAAI 2024technical

Federated Graph Learning is a privacy-preserving collaborative approach for training a shared model on graph-structured data in the distributed environment. However, in real-world scenarios, the client graph data usually originate from diverse domains, this unavoidably hinders the generalization per…

2024

Parameter Disparities Dissection for Backdoor Defense in Heterogeneous Federated Learning

NeurIPS 2024poster

Backdoor attacks pose a serious threat to federated systems, where malicious clients optimize on the triggered distribution to mislead the global model towards a predefined target. Existing backdoor defense methods typically require either homogeneous assumption, validation datasets, or client optim…

Cited by 3SourcePDFScholar
2024

S3GCL: Spectral, Swift, Spatial Graph Contrastive Learning

ICML 2024poster

Graph Contrastive Learning (GCL) has emerged as a highly effective self-supervised approach in graph representation learning. However, prevailing GCL methods confront two primary challenges: 1) They predominantly operate under homophily assumptions, focusing on low-frequency signals in node features…

Cited by 15SourcePDFScholar
2024

Self-Driven Entropy Aggregation for Byzantine-Robust Heterogeneous Federated Learning

ICML 2024poster

Federated learning presents massive potential for privacy-friendly collaboration. However, the performance of federated learning is deeply affected by byzantine attacks, where malicious clients deliberately upload crafted vicious updates. While various robust aggregations have been proposed to defen…

Cited by 5SourcePDFScholar
2024

Shallow-Deep Collaborative Learning for Unsupervised Visible-Infrared Person Re-Identification

CVPR 2024poster

Unsupervised visible-infrared person re-identification (US-VI-ReID) centers on learning a cross-modality retrieval model without labels reducing the reliance on expensive cross-modality manual annotation. Previous US-VI-ReID works gravitate toward learning cross-modality information with the deep fe…

2023

Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval

CVPR 2023poster

Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have attempted to address this challenge by leveraging separately pre-t…

2023

Dynamic Personalized Federated Learning with Adaptive Differential Privacy

NeurIPS 2023poster

Personalized federated learning with differential privacy has been considered a feasible solution to address non-IID distribution of data and privacy leakage risks. However, current personalized federated learning methods suffer from inflexible personalization and convergence difficulties due to two…

2023

Prototype Reminiscence and Augmented Asymmetric Knowledge Aggregation for Non-Exemplar Class-Incremental Learning

ICCV 2023poster

Non-exemplar class-incremental learning (NECIL) requires deep models to maintain existing knowledge while continuously learning new classes without saving old class samples. In NECIL methods, prototypical representations are usually stored, which inject information from former classes to resist cata…

Cited by 42PDFScholar
2023

Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning

AAAI 2023technical

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at…

2023

Rethinking Federated Learning With Domain Shift: A Prototype View

CVPR 2023poster

Federated learning shows a bright promise as a privacy-preserving collaborative learning technique. However, prevalent solutions mainly focus on all private data sampled from the same domain. An important challenge is that when distributed data are derived from diverse domains. The private model pre…

2023

Top-K Visual Tokens Transformer: Selecting Tokens for Visible-Infrared Person Re-Identification

ICASSP 2023accepted

Visible modality and infrared modality person re-identification (VI-ReID) is an extremely important and challenging task. Existing works mainly focus on reducing the modality gap with Convolutional Neural Networks (CNN). However, the features extracted by CNN may contain useless identity-irrelevant…

Cited by 0SourceScholar
2023

Towards Grand Unified Representation Learning for Unsupervised Visible-Infrared Person Re-Identification

ICCV 2023poster

Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) is an extremely important and challenging task, which can alleviate the issue of expensive cross-modality annotations. Existing works focus on handling the cross-modality discrepancy under unsupervised conditions. However,…

Cited by 52PDFcodeScholar
2023

Towards Modality-Agnostic Person Re-Identification With Descriptive Query

CVPR 2023poster

Person re-identification (ReID) with descriptive query (text or sketch) provides an important supplement for general image-image paradigms, which is usually studied in a single cross-modality matching manner, e.g., text-to-image or sketch-to-photo. However, without a camera-captured photo query, it…

2023

Unsupervised Visible-Infrared Person Re-Identification via Progressive Graph Matching and Alternate Learning

CVPR 2023poster

Unsupervised visible-infrared person re-identification is a challenging task due to the large modality gap and the unavailability of cross-modality correspondences. Cross-modality correspondences are very crucial to bridge the modality gap. Some existing works try to mine cross-modality corresponden…

2021

Cooperative Joint Attentive Network for Patient Outcome Prediction on Irregular Multi-Rate Multivariate Health Data

IJCAI 2021poster

Due to the dynamic health status of patients and discrepant stability of physiological variables, health data often presents as irregular multi-rate multivariate time series (IMR-MTS) with significantly varying sampling rates. Existing methods mainly study changes of IMR-MTS values in the time domai…

Cited by 11SourcePDFScholar
2021

Cross-Modality Person Re-Identification via Modality Confusion and Center Aggregation

ICCV 2021poster

Cross-modality person re-identification is a challenging task due to large cross-modality discrepancy and intra-modality variations. Currently, most existing methods focus on learning modality-specific or modality-shareable features by using the identity supervision or modality label. Different from…

Cited by 213PDFScholar
2020

Dynamic Dual-Attentive Aggregation Learning for Visible-Infrared Person Re-Identification

ECCV 2020poster

Visible-infrared person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval problem. Due to the large intra-class variations and cross-modality discrepancy with large amount of sample noise, it is difficult to learn discriminative part features. Existing VI-ReID methods…

2019

Unsupervised Embedding Learning via Invariant and Spreading Instance Feature

CVPR 2019poster

This paper studies the unsupervised embedding learning problem, which requires an effective similarity measurement between samples in low-dimensional embedding space. Motivated by the positive concentrated and negative separated properties observed from category-wise supervised learning, we propose…

Cited by 760PDFcodeScholar
2018

Robust Anchor Embedding for Unsupervised Video Person Re-Identification in the Wild

ECCV 2018poster

This paper addresses the scalability and robustness issues of estimating labels from imbalanced unlabeled data for unsupervised video-based person re-identification (re-ID). To achieve it, we propose a novel Robust AnChor Embedding (RACE) framework via deep feature representation learning for large-…

Cited by 127SourcePDFScholar
2017

Dynamic Label Graph Matching for Unsupervised Video Re-Identification

ICCV 2017poster

Label estimation is an important component in an unsupervised person re-identification (re-ID) system. This paper focuses on cross-camera label estimation, which can be subsequently used in feature learning to learn robust re-ID models. Specifically, we propose to construct a graph for samples in ea…

Cited by 224PDFScholar