← Search

Dezhong Peng

41 accepted papers

2026

Ambiguity-Tolerant Cross-Modal Hashing with Partial Labels

AAAI 2026technical

Cross-modal hashing (CMH) has achieved remarkable success in large-scale cross-modal retrieval due to its low storage cost and high computational efficiency. However, most existing CMH methods rely on accurately annotated training data, which is often impractical in real-world applications due to th

Cited by 0SourcePDFScholar
2026

Correspondence Cognitive Learning for Multi-Modal Object Re-Identification

ICML 2026poster

Multi-modal object Re-Identification (ReID) aims to retrieve the same object across different modalities by exploiting their complementary visual information. Recent advances leverage Multi-modal Large Language Models (MLLMs) to generate descriptive textual annotations as auxiliary supervision. Howe…

Cited by 0SourceScholar
2026

Learning Beyond Domains: Misleading Prompts and Pseudo-Label Contrast for Text Domain Generalization

AAAI 2026technical

Recent advancements in Pre-trained Language Models (PLMs) have significantly enhanced performance across various Natural Language Processing (NLP) tasks. However, the variability in data distributions across different domains presents challenges in generalizing these models to unseen domains. Domain

Cited by 0SourcePDFScholar
2026

Multimodal Nested Learning for Decoupled and Coordinated Optimization

ICML 2026oral

Multimodal learning aims to integrate multi-sensor data to exploit their complementary information, embracing a more comprehensive real-world perception and understanding. However, heterogeneous discrepancies across modalities consistently trigger imbalanced multimodal optimization, restricting the …

Cited by 0SourceScholar
2026

Robust Semi-paired Multimodal Learning for Cross-modal Retrieval

AAAI 2026technical

Cross-modal retrieval is a fundamental application of multi-modal learning that has achieved remarkable success with large-scale well-paired data. However, in practice, it is costly to collect large-scale well-paired data. To alleviate the dependence on the amount of paired data, in this paper, we s

Cited by 0SourcePDFScholar
2026

Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval

AAAI 2026technical

Cross-modal hashing (CMH) facilitates efficient retrieval across different modalities (e.g., image and text) by encoding data into compact binary representations. While recent methods have achieved remarkable performance, they often rely heavily on fully annotated datasets, which are costly and labo

Cited by 0SourcePDFScholar
2025

CoPINN: Cognitive Physics-Informed Neural Networks

ICML 2025spotlight

Physics-informed neural networks (PINNs) aim to constrain the outputs and gradients of deep learning models to satisfy specified governing physics equations, which have demonstrated significant potential for solving partial differential equations (PDEs). Although existing PINN methods have achieved…

Cited by 0SourcePDFScholar
2025

Deep Evidential Hashing for Trustworthy Cross-Modal Retrieval

AAAI 2025technical

Cross-modal hashing provides an efficient solution for retrieval tasks across various modalities, such as images and text. However, most existing methods are deterministic models, which overlook the reliability associated with the retrieved results. This omission renders them unreliable for determin…

2025

Deep Fuzzy Multi-view Learning for Reliable Classification

ICML 2025poster

Multi-view learning methods primarily focus on enhancing decision accuracy but often neglect the uncertainty arising from the intrinsic drawbacks of data, such as noise, conflicts, etc. To address this issue, several trusted multi-view learning approaches based on the Evidential Theory have been pro…

Cited by 0SourcePDFScholar
2025

DiCA: Disambiguated Contrastive Alignment for Cross-Modal Retrieval with Partial Labels

AAAI 2025technical

Cross-modal retrieval aims to retrieve relevant data across different modalities. Driven by costly massive labeled data, existing cross-modal retrieval methods achieve encouraging results. To reduce annotation costs while maintaining performance, this paper focuses on an untouched but challenging pr…

2025

Fuzzy Multimodal Learning for Trusted Cross-modal Retrieval

CVPR 2025poster

Cross-modal retrieval aims to match related samples across distinct modalities, facilitating the retrieval and discovery of heterogeneous information. Although existing methods show promising performance, most are deterministic models and are unable to capture the uncertainty inherent in the retriev…

2025

Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification

CVPR 2025poster

Despite remarkable advancements in text-to-image person re-identification (TIReID) facilitated by the breakthrough of cross-modal embedding models, existing methods often struggle to distinguish challenging candidate images due to intrinsic limitations, such as network architecture and data quality.…

2025

Interactive Cross-modal Learning for Text-3D Scene Retrieval

NeurIPS 2025oral

Text-3D Scene Retrieval (T3SR) aims to retrieve relevant scenes using linguistic queries. Although traditional T3SR methods have made significant progress in capturing fine-grained associations, they implicitly assume that query descriptions are information-complete. In practical deployments, howeve…

Cited by 0SourceScholar
2025

LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification

ICML 2025poster

Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often partial or vague. To address this limitation, we introduce a new task called interactive person re-identification (Inter…

2025

Learning Source-Free Domain Adaptation for Visible-Infrared Person Re-Identification

NeurIPS 2025poster

In this paper, we investigate source-free domain adaptation (SFDA) for visible-infrared person re-identification (VI-ReID), aiming to adapt a pre-trained source model to an unlabeled target domain without access to source data. To address this challenging setting, we propose a novel learning paradig…

Cited by 0SourceScholar
2025

Noisy Label Calibration for Multi-View Classification

AAAI 2025technical

In recent years, multi-view learning has aroused extensive research passion. Most existing multi-view learning methods often rely on well-annotations to improve decision accuracy. However, noise labels are ubiquitous in multi-view data due to imperfect annotations. To deal with this problem, we prop…

2025

ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy Correspondence

CVPR 2025highlight

Multi-view clustering (MVC) aims to exploit complementary information from diverse views to enhance clustering performance. Since pseudo-labels can provide additional semantic information, many MVC methods have been proposed to guide unsupervised multi-view learning through pseudo-labels. These meth…

Cited by 0SourcePDFScholar
2025

ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL

ICLR 2025poster

Despite the significant advancements in Text-to-SQL (Text2SQL) facilitated by large language models (LLMs), the latest state-of-the-art techniques are still trapped in the in-context learning of closed-source LLMs (e.g., GPT-4), which limits their applicability in open scenarios. To address this ch…

2025

Reliable Disentanglement Multi-view Learning Against View Adversarial Attacks

IJCAI 2025

Trustworthy multi-view learning has attracted extensive attention because evidence learning can provide reliable uncertainty estimation to enhance the credibility of multi-view predictions. Existing trusted multi-view learning methods implicitly assume that multi-view data is secure. However, in saf

2025

RoDA: Robust Domain Alignment for Cross-Domain Retrieval Against Label Noise

AAAI 2025technical

This paper studies the complex challenge of cross-domain image retrieval under the condition of noisy labels (NCIR), a scenario that not only includes the inherent obstacles of traditional cross-domain image retrieval (CIR) but also requires alleviating the adverse effects of label noise. To address…

2025

Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and Grounding

NeurIPS 2025poster

Grounding target objects in 3D environments via natural language is a fundamental capability for autonomous agents to successfully fulfill user requests. Almost all existing works typically assume that the target object lies within a known scene and focus solely on in-scene localization. In practice…

Cited by 0SourceScholar
2025

Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels

AAAI 2025technical

Cross-modal hashing (CMH) has appeared as a popular technique for cross-modal retrieval due to its low storage cost and high computational efficiency in large-scale data. Most existing methods implicitly assume that multi-modal data is correctly labeled, which is expensive and even unattainable due…

2024

Decoupled Contrastive Multi-View Clustering with High-Order Random Walks

AAAI 2024technical

In recent, some robust contrastive multi-view clustering (MvC) methods have been proposed, which construct data pairs from neighborhoods to alleviate the false negative issue, i.e., some intra-cluster samples are wrongly treated as negative pairs. Although promising performance has been achieved by…

2024

DiDA: Disambiguated Domain Alignment for Cross-Domain Retrieval with Partial Labels

AAAI 2024technical

Driven by generative AI and the Internet, there is an increasing availability of a wide variety of images, leading to the significant and popular task of cross-domain image retrieval. To reduce annotation costs and increase performance, this paper focuses on an untouched but challenging problem, i.e…

2024

Dual Semantic Fusion Hashing for Multi-Label Cross-Modal Retrieval

IJCAI 2024poster

Cross-modal hashing (CMH) has been widely used for multi-modal retrieval tasks due to its low storage cost and fast query speed. Although existing CMH methods achieve promising performance, most of them mainly rely on coarse-grained supervision information (\ie pairwise similarity matrix) to measure…

Cited by 4SourcePDFScholar
2024

Image Clustering with External Guidance

ICML 2024oral

The core of clustering lies in incorporating prior knowledge to construct supervision signals. From classic k-means based on data compactness to recent contrastive clustering guided by self-supervision, the evolution of clustering methods intrinsically corresponds to the progression of supervision s…

2024

Noisy-Correspondence Learning for Text-to-Image Person Re-identification

CVPR 2024poster

Text-to-image person re-identification (TIReID) is a compelling topic in the cross-modal community which aims to retrieve the target person based on a textual query. Although numerous TIReID methods have been proposed and achieved promising performance they implicitly assume the training image-text…

2023

Comprehensive and Delicate: An Efficient Transformer for Image Restoration

CVPR 2023poster

Vision Transformers have shown promising performance in image restoration, which usually conduct window- or channel-based attention to avoid intensive computations. Although the promising performance has been achieved, they go against the biggest success factor of Transformers to a certain extent by…

2023

Correspondence-Free Domain Alignment for Unsupervised Cross-Domain Image Retrieval

AAAI 2023technical

Cross-domain image retrieval aims at retrieving images across different domains to excavate cross-domain classificatory or correspondence relationships. This paper studies a less-touched problem of cross-domain image retrieval, i.e., unsupervised cross-domain image retrieval, considering the followi…

2023

Cross-modal Active Complementary Learning with Self-refining Correspondence

NeurIPS 2023poster

Recently, image-text matching has attracted more and more attention from academia and industry, which is fundamental to understanding the latent correspondence across visual and textual modalities. However, most existing methods implicitly assume the training pairs are well-aligned while ignoring th…

2023

Deep Fair Clustering via Maximizing and Minimizing Mutual Information: Theory, Algorithm and Metric

CVPR 2023poster

Fair clustering aims to divide data into distinct clusters while preventing sensitive attributes (e.g., gender, race, RNA sequencing technique) from dominating the clustering. Although a number of works have been conducted and achieved huge success recently, most of them are heuristical, and there l…

2023

Dynamic Voting for Efficient Reasoning in Large Language Models

EMNLP 2023long findings

Multi-path voting methods like Self-consistency have been used to mitigate reasoning errors in large language models caused by factual errors and illusion generation. However, these methods require excessive computing resources as they generate numerous reasoning paths for each problem. And our expe…

Cited by 0SourceScholar
2023

Incomplete Multi-view Clustering via Prototype-based Imputation

IJCAI 2023poster

In this paper, we study how to achieve two characteristics highly-expected by incomplete multi-view clustering (IMvC). Namely, i) instance commonality refers to that within-cluster instances should share a common pattern, and ii) view versatility refers to that cross-view samples should own view-spe…

2023

RONO: Robust Discriminative Learning With Noisy Labels for 2D-3D Cross-Modal Retrieval

CVPR 2023poster

Recently, with the advent of Metaverse and AI Generated Content, cross-modal retrieval becomes popular with a burst of 2D and 3D data. However, this problem is challenging given the heterogeneous structure and semantic discrepancies. Moreover, imperfect annotations are ubiquitous given the ambiguous…

2023

Unifying Discrete and Continuous Representations for Unsupervised Paraphrase Generation

EMNLP 2023long main

Unsupervised paraphrase generation is a challenging task that benefits a variety of downstream NLP applications. Current unsupervised methods for paraphrase generation typically employ round-trip translation or denoising, which require translation corpus and result in paraphrases overly similar to t…

Cited by 0SourceScholar
2022

Adaptive Meta-learner via Gradient Similarity for Few-shot Text Classification

COLING 2022main

Few-shot text classification aims to classify the text under the few-shot scenario. Most of the previous methods adopt optimization-based meta learning to obtain task distribution. However, due to the neglect of matching between the few amount of samples and complicated models, as well as the distin…