← Search

Peng Hu

66 accepted papers

2026

Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image Restoration

CVPR 2026

Real-world image restoration aims to restore high-quality (HQ) images from degraded low-quality (LQ) inputs captured under uncontrolled conditions. Existing methods typically depend on ground-truth (GT) supervision, assuming that GT provides perfect reference quality. However, GT can still contain i

Cited by 4SourcecodeScholar
2026

Bootstrapping Multi-view Learning for Test-time Noisy Correspondence

CVPR 2026

Multi-view learning fuses complementary views to improve perception, but real-world deployments often suffer from Test-time Noisy Correspondence (TNC) -- cross-view misalignment caused by asynchronous sampling, transient network congestion, or other disturbances. Such misalignment introduces semanti

Cited by 0SourcecodeScholar
2026

DOUBT: Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness for Hallucination Detection in MLLMs

ICML 2026oral

Multimodal Large Language Models (MLLMs) frequently produce hallucinations (i.e., assertions that contradict the image or facts), undermining reliability in high-risk applications. Existing detection approaches typically feed images and texts jointly and estimate hallucination scores by measuring th…

Cited by 0SourceScholar
2026

Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy Labels

ICML 2026spotlight

Recently, cross-modal hashing (CMH) has garnered significant attention due to its low storage costs and high retrieval efficiency. most existing CMH methods implicitly assume the availability of high-quality annotations, which is often violated in real-world scenarios as label noise inevitably arise…

Cited by 0SourceScholar
2026

Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment

ICLR 2026oral

Multi-modal entity alignment (MMEA) aims to identify equivalent entities across heterogeneous multi-modal knowledge graphs (MMKGs), where each entity is described by attributes from various modalities. Existing methods typically assume that both intra-entity and inter-graph correspondences are fault…

Cited by 0SourcecodeScholar
2026

Multimodal Nested Learning for Decoupled and Coordinated Optimization

ICML 2026oral

Multimodal learning aims to integrate multi-sensor data to exploit their complementary information, embracing a more comprehensive real-world perception and understanding. However, heterogeneous discrepancies across modalities consistently trigger imbalanced multimodal optimization, restricting the …

Cited by 0SourceScholar
2026

Next-Scale Prediction: A Self-Supervised Approach for Real-World Image Denoising

CVPR 2026

Self-supervised real-world image denoising remains a fundamental challenge, arising from the antagonistic trade-off between decorrelating spatially structured noise and preserving high-frequency details. Existing blind-spot network (BSN) methods rely on pixel-shuffle downsampling (PD) to decorrelate

Cited by 0SourcecodeScholar
2026

RLSF-V: Mitigating Hallucinations in MLLMs via Fuzzy Semantic Self-Feedback

ICML 2026poster

Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception for open-world understanding, but exacerbate LLMs' hallucinations, in which generated text contradicts visual evidence or common sense. To mitigate hallucinations, a dominant strategy is Direct Prefer…

Cited by 0SourceScholar
2026

Robust Semi-paired Multimodal Learning for Cross-modal Retrieval

AAAI 2026technical

Cross-modal retrieval is a fundamental application of multi-modal learning that has achieved remarkable success with large-scale well-paired data. However, in practice, it is costly to collect large-scale well-paired data. To alleviate the dependence on the amount of paired data, in this paper, we s

Cited by 0SourcePDFScholar
2026

Uncover Underlying Correspondence for Robust Multi-view Clustering

ICLR 2026oral

Multi-view clustering (MVC) aims to group unlabeled data into semantically meaningful clusters by leveraging cross-view consistency. However, real-world datasets collected from the web often suffer from noisy correspondence (NC), which breaks the consistency prior and results in unreliable alignmen…

Cited by 0SourcecodeScholar
2025

CoPINN: Cognitive Physics-Informed Neural Networks

ICML 2025spotlight

Physics-informed neural networks (PINNs) aim to constrain the outputs and gradients of deep learning models to satisfy specified governing physics equations, which have demonstrated significant potential for solving partial differential equations (PDEs). Although existing PINN methods have achieved…

Cited by 0SourcePDFScholar
2025

Conditional Representation Learning for Customized Tasks

NeurIPS 2025spotlight

Conventional representation learning methods learn a universal representation that primarily captures dominant semantics, which may not always align with customized downstream tasks. For instance, in animal habitat analysis, researchers prioritize scene-related features, whereas universal embeddings…

Cited by 0SourcecodeScholar
2025

Deep Evidential Hashing for Trustworthy Cross-Modal Retrieval

AAAI 2025technical

Cross-modal hashing provides an efficient solution for retrieval tasks across various modalities, such as images and text. However, most existing methods are deterministic models, which overlook the reliability associated with the retrieved results. This omission renders them unreliable for determin…

2025

Deep Fuzzy Multi-view Learning for Reliable Classification

ICML 2025poster

Multi-view learning methods primarily focus on enhancing decision accuracy but often neglect the uncertainty arising from the intrinsic drawbacks of data, such as noise, conflicts, etc. To address this issue, several trusted multi-view learning approaches based on the Evidential Theory have been pro…

Cited by 0SourcePDFScholar
2025

Deep Unsupervised Hashing via External Guidance

ICML 2025poster

Recently, deep unsupervised hashing has gained considerable attention in image retrieval due to its advantages in cost-free data labeling, computational efficiency, and storage savings. Although existing methods achieve promising performance by leveraging inherent visual structures within the data,…

Cited by 0SourcePDFScholar
2025

Disentangling Multi-view Representations via Curriculum Learning with Learnable Prior

IJCAI 2025

Multi-view representation learning methods typically follow a consistent-and-specific pipeline that aims at extracting latent representations for an entity from its multiple observable views to facilitate downstream tasks. However, most of them overlook the complex underlying correlation between dif

2025

Fuzzy Multimodal Learning for Trusted Cross-modal Retrieval

CVPR 2025poster

Cross-modal retrieval aims to match related samples across distinct modalities, facilitating the retrieval and discovery of heterogeneous information. Although existing methods show promising performance, most are deterministic models and are unable to capture the uncertainty inherent in the retriev…

2025

Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification

CVPR 2025poster

Despite remarkable advancements in text-to-image person re-identification (TIReID) facilitated by the breakthrough of cross-modal embedding models, existing methods often struggle to distinguish challenging candidate images due to intrinsic limitations, such as network architecture and data quality.…

2025

Interactive Cross-modal Learning for Text-3D Scene Retrieval

NeurIPS 2025oral

Text-3D Scene Retrieval (T3SR) aims to retrieve relevant scenes using linguistic queries. Although traditional T3SR methods have made significant progress in capturing fine-grained associations, they implicitly assume that query descriptions are information-complete. In practical deployments, howeve…

Cited by 0SourceScholar
2025

Known-Plaintext Attacks to Thumbnail-Preservation Encryption Using Pix2pix Generative Adversarial Network

ICASSP 2025accepted

General image encryption schemes transform plaintext images into snowflake-like ciphertext patterns through efficient cryptographic permutation and confusion primitives. However, these encryption schemes cannot achieve the goal of privacy computing that data are available but not visible. Thumbnail-…

Cited by 0SourceScholar
2025

LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification

ICML 2025poster

Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often partial or vague. To address this limitation, we introduce a new task called interactive person re-identification (Inter…

2025

Large Language Models Are Cross-Lingual Knowledge-Free Reasoners

NAACL 2025long

Large Language Models have demonstrated impressive reasoning capabilities across multiple languages. However, the relationship between capabilities in different languages is less explored. In this work, we decompose the process of reasoning tasks into two separated components: knowledge retrieval an…

2025

Learning Robust Multi-view Representation Using Dual-masked VAEs

IJCAI 2025

Most existing multi-view representation learning methods assume view-completeness and noise-free data. However, such assumptions are strong in real-world applications. Despite advances in methods tailored to view-missing or noise problems individually, a one-size-fits-all approach that concurrently

2025

Learning Source-Free Domain Adaptation for Visible-Infrared Person Re-Identification

NeurIPS 2025poster

In this paper, we investigate source-free domain adaptation (SFDA) for visible-infrared person re-identification (VI-ReID), aiming to adapt a pre-trained source model to an unlabeled target domain without access to source data. To address this challenging setting, we propose a novel learning paradig…

Cited by 0SourceScholar
2025

Learning with Noisy Triplet Correspondence for Composed Image Retrieval

CVPR 2025poster

Composed Image Retrieval (CIR) enables editable image search by integrating a query pair--a reference image ref and a textual modification mod--to retrieve a target image tar that reflects the intended change. While existing CIR methods have shown promising performance using well-annotated triplets…

2025

MaIR: A Locality- and Continuity-Preserving Mamba for Image Restoration

CVPR 2025poster

Recent advancements in Mamba have shown promising results in image restoration. These methods typically flatten 2D images into multiple distinct 1D sequences along rows and columns, process each sequence independently using selective scan operation, and recombine them to form the outputs. However, s…

2025

Probabilistic Multimodal Learning with von Mises-Fisher Distributions

IJCAI 2025

Multimodal learning is pivotal for the advancement of artificial intelligence, enabling machines to integrate complementary information from diverse data sources for holistic perception and understanding. Despite significant progress, existing methods struggle with challenges such as noisy inputs, n

2025

ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy Correspondence

CVPR 2025highlight

Multi-view clustering (MVC) aims to exploit complementary information from diverse views to enhance clustering performance. Since pseudo-labels can provide additional semantic information, many MVC methods have been proposed to guide unsupervised multi-view learning through pseudo-labels. These meth…

Cited by 0SourcePDFScholar
2025

ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL

ICLR 2025poster

Despite the significant advancements in Text-to-SQL (Text2SQL) facilitated by large language models (LLMs), the latest state-of-the-art techniques are still trapped in the in-context learning of closed-source LLMs (e.g., GPT-4), which limits their applicability in open scenarios. To address this ch…

2025

Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and Grounding

NeurIPS 2025poster

Grounding target objects in 3D environments via natural language is a fundamental capability for autonomous agents to successfully fulfill user requests. Almost all existing works typically assume that the target object lies within a known scene and focus solely on in-scene localization. In practice…

Cited by 0SourceScholar
2025

Test-time Adaptation for Cross-modal Retrieval with Query Shift

ICLR 2025spotlight

The success of most existing cross-modal retrieval methods heavily relies on the assumption that the given queries follow the same distribution of the source domain. However, such an assumption is easily violated in real-world scenarios due to the complexity and diversity of queries, thus leading t…

2025

Visual Abstraction: A Plug-and-Play Approach for Text-Visual Retrieval

ICML 2025poster

Text-to-visual retrieval often struggles with semantic redundancy and granularity mismatches between textual queries and visual content. Unlike existing methods that address these challenges during training, we propose VISual Abstraction (VISA), a test-time approach that enhances retrieval by transf…

Cited by 0SourcePDFScholar
2024

AverNet: All-in-one Video Restoration for Time-varying Unknown Degradations

NeurIPS 2024poster

Traditional video restoration approaches were designed to recover clean videos from a specific type of degradation, making them ineffective in handling multiple unknown types of degradation. To address this issue, several studies have been conducted and have shown promising results. However, these s…

2024

Decoupled Contrastive Multi-View Clustering with High-Order Random Walks

AAAI 2024technical

In recent, some robust contrastive multi-view clustering (MvC) methods have been proposed, which construct data pairs from neighborhoods to alleviate the false negative issue, i.e., some intra-cluster samples are wrongly treated as negative pairs. Although promising performance has been achieved by…

2024

Image Clustering with External Guidance

ICML 2024oral

The core of clustering lies in incorporating prior knowledge to construct supervision signals. From classic k-means based on data compactness to recent contrastive clustering guided by self-supervision, the evolution of clustering methods intrinsically corresponds to the progression of supervision s…

2024

Large Language Models are Limited in Out-of-Context Knowledge Reasoning

EMNLP 2024finding

Large Language Models (LLMs) possess extensive knowledge and strong capabilities in performing in-context reasoning. However, previous work challenges their out-of-context reasoning ability, i.e., the ability to infer information from their training data, instead of from the context or prompt. This…

2024

Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only Shallowly

NAACL 2024long

Despite their strong ability to retrieve knowledge in English, current large language models show imbalance abilities in different languages. Two approaches are proposed to address this, i.e., multilingual pretraining and multilingual instruction tuning. However, whether and how do such methods cont…

2024

Noisy-Correspondence Learning for Text-to-Image Person Re-identification

CVPR 2024poster

Text-to-image person re-identification (TIReID) is a compelling topic in the cross-modal community which aims to retrieve the target person based on a textual query. Although numerous TIReID methods have been proposed and achieved promising performance they implicitly assume the training image-text…

2024

PrefAce: Face-Centric Pretraining with Self-Structure Aware Distillation

AAAI 2024technical

Video-based facial analysis is important for autonomous agents to understand human expressions and sentiments. However, limited labeled data is available to learn effective facial representations. This paper proposes a novel self-supervised face-centric pretraining framework, called PrefAce, which l…

2024

Robust Contrastive Multi-view Clustering against Dual Noisy Correspondence

NeurIPS 2024poster

Recently, contrastive multi-view clustering (MvC) has emerged as a promising avenue for analyzing data from heterogeneous sources, typically leveraging the off-the-shelf instances as positives and randomly sampled ones as negatives. In practice, however, this paradigm would unavoidably suffer from t…

2024

Test-time Adaptation against Multi-modal Reliability Bias

ICLR 2024poster

Test-time adaptation (TTA) has emerged as a new paradigm for reconciling distribution shifts across domains without accessing source data. However, existing TTA methods mainly concentrate on uni-modal tasks, overlooking the complexity of multi-modal scenarios. In this paper, we delve into the multi-…

2023

Correspondence-Free Domain Alignment for Unsupervised Cross-Domain Image Retrieval

AAAI 2023technical

Cross-domain image retrieval aims at retrieving images across different domains to excavate cross-domain classificatory or correspondence relationships. This paper studies a less-touched problem of cross-domain image retrieval, i.e., unsupervised cross-domain image retrieval, considering the followi…

2023

Cross-modal Active Complementary Learning with Self-refining Correspondence

NeurIPS 2023poster

Recently, image-text matching has attracted more and more attention from academia and industry, which is fundamental to understanding the latent correspondence across visual and textual modalities. However, most existing methods implicitly assume the training pairs are well-aligned while ignoring th…

2023

Deep Fair Clustering via Maximizing and Minimizing Mutual Information: Theory, Algorithm and Metric

CVPR 2023poster

Fair clustering aims to divide data into distinct clusters while preventing sensitive attributes (e.g., gender, race, RNA sequencing technique) from dominating the clustering. Although a number of works have been conducted and achieved huge success recently, most of them are heuristical, and there l…

2023

Graph Matching with Bi-level Noisy Correspondence

ICCV 2023poster

In this paper, we study a novel and widely existing problem in graph matching (GM), namely, Bi-level Noisy Correspondence (BNC), which refers to node-level noisy correspondence (NNC) and edge-level noisy correspondence (ENC). In brief, on the one hand, due to the poor recognizability and viewpoint d…

Cited by 42PDFcodeScholar
2023

Incomplete Multi-view Clustering via Prototype-based Imputation

IJCAI 2023poster

In this paper, we study how to achieve two characteristics highly-expected by incomplete multi-view clustering (IMvC). Namely, i) instance commonality refers to that within-cluster instances should share a common pattern, and ii) view versatility refers to that cross-view samples should own view-spe…

2023

RONO: Robust Discriminative Learning With Noisy Labels for 2D-3D Cross-Modal Retrieval

CVPR 2023poster

Recently, with the advent of Metaverse and AI Generated Content, cross-modal retrieval becomes popular with a burst of 2D and 3D data. However, this problem is challenging given the heterogeneous structure and semantic discrepancies. Moreover, imperfect annotations are ubiquitous given the ambiguous…

2023

Rethinking Image Super Resolution From Long-Tailed Distribution Learning Perspective

CVPR 2023poster

Existing studies have empirically observed that the resolution of the low-frequency region is easier to enhance than that of the high-frequency one. Although plentiful works have been devoted to alleviating this problem, little understanding is given to explain it. In this paper, we try to give a fe…

Cited by 15SourcePDFScholar
2022

All-in-One Image Restoration for Unknown Corruption

CVPR 2022poster

In this paper, we study a challenging problem in image restoration, namely, how to develop an all-in-one method that could recover images from a variety of unknown corruption types and levels. To this end, we propose an All-in-one Image Restoration Network (AirNet) consisting of two neural modules,…

Cited by 351PDFcodeScholar
2022

Learning With Twin Noisy Labels for Visible-Infrared Person Re-Identification

CVPR 2022poster

In this paper, we study an untouched problem in visible-infrared person re-identification (VI-ReID), namely, Twin Noise Labels (TNL) which refers to as noisy annotation and correspondence. In brief, on the one hand, it is inevitable to annotate some persons with the wrong identity due to the complex…

Cited by 213PDFcodeScholar
2022

Multi-Scale Adaptive Network for Single Image Denoising

NeurIPS 2022accept

Multi-scale architectures have shown effectiveness in a variety of tasks thanks to appealing cross-scale complementarity. However, existing architectures treat different scale features equally without considering the scale-specific characteristics, \textit{i.e.}, the within-scale characteristics are…

2021

A Primal-Dual Online Algorithm for Online Matching Problem in Dynamic Environments

AAAI 2021technical

Recently, the online matching problem has attracted much attention due to its wide application on real-world decision-making scenarios. In stationary environments, by adopting the stochastic user arrival model, existing methods are proposed to learn dual optimal prices and are shown to achieve a fas…

Cited by 2SourcePDFScholar
2021

BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction

ICLR 2021poster

We study the challenging task of neural network quantization without end-to-end retraining, called Post-training Quantization (PTQ). PTQ usually requires a small subset of training data but produces less powerful quantized models than Quantization-Aware Training (QAT). In this work, we propose a nov…

2021

Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis

ICASSP 2021accepted

LiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from…

Cited by 0SourceScholar
2021

OPQ: Compressing Deep Neural Networks with One-shot Pruning-Quantization

AAAI 2021technical

As Deep Neural Networks (DNNs) usually are overparameterized and have millions of weight parameters, it is challenging to deploy these large DNN models on resource-constrained hardware platforms, e.g., smartphones. Numerous network compression methods such as pruning and quantization are proposed to…

Cited by 69SourcePDFScholar
2021

PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS With Relationship Recovery

CVPR 2021poster

Non-maximum Suppression (NMS) is an essential post-processing step in modern convolutional neural networks for object detection. Unlike convolutions which are inherently parallel, the de-facto standard for NMS, namely GreedyNMS, cannot be easily parallelized and thus could be the performance bottlen…

Cited by 12PDFcodeScholar
2021

Partially View-Aligned Representation Learning With Noise-Robust Contrastive Loss

CVPR 2021poster

In real-world applications, it is common that only a portion of data is aligned across views due to spatial, temporal, or spatiotemporal asynchronism, thus leading to the so-called Partially View-aligned Problem (PVP). To solve such a less-touched problem without the help of labels, we propose simul…

Cited by 170PDFScholar
2019

Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks

ICCV 2019poster

Hardware-friendly network quantization (e.g., binary/uniform quantization) can efficiently accelerate the inference and meanwhile reduce memory consumption of the deep neural networks, which is crucial for model deployment on resource-limited devices like mobile phones. However, due to the discreten…

Cited by 591PDFScholar