← Search

Lihe Zhang

28 accepted papers

2026

Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection

CVPR 2026

Multimodal unsupervised anomaly detection has garnered increasing attention for robust defect localization.Recent approaches rely on establishing cross-modal matching relationships under normal conditions without explicit guidance.However, in practice, a single modality may have multiple distinct re

Cited by 0SourcecodeScholar
2026

Mitigating Error Amplification in Fast Adversarial Training

CVPR 2026

Fast Adversarial Training (FAT) has proven effective in enhancing model robustness by encouraging networks to learn perturbation-invariant representations.However, FAT often suffers from catastrophic overfitting (CO), where the model overfits to the training attack and fails to generalize to unseen

Cited by 0SourceScholar
2026

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

AAAI 2026technical

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during pa

Cited by 0SourcePDFScholar
2026

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

CVPR 2026

Existing anomaly detection methods often treat the modality and class as independent factors. Although this paradigm has enriched the development of AD research branches and produced many specialized models, it has also led to fragmented solutions and excessive memory overhead. Moreover, reconstruct

Cited by 0SourcecodeScholar
2025

High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity

ICLR 2025poster

In the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets co…

2025

Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation

CVPR 2025poster

In this paper, we present a new dynamic collaborative network for semi-supervised 3D vessel segmentation, termed DiCo. Conventional mean teacher (MT) methods typically employ a static approach, where the roles of the teacher and student models are fixed. However, due to the complexity of 3D vessel d…

2025

Rethinking Evaluation of Infrared Small Target Detection

NeurIPS 2025poster

As an essential vision task, infrared small target detection (IRSTD) has seen significant advancements through deep learning. However, critical limitations in current evaluation protocols impede further progress. First, existing methods rely on fragmented pixel- and target-level specific met…

Cited by 0SourceScholar
2025

UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation

NeurIPS 2025poster

Multi-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they introduce high deployment costs by requiring exhaustive mode…

Cited by 0SourcecodeScholar
2025

Unified Medical Lesion Segmentation via Self-referring Indicator

CVPR 2025poster

The recently emerged in-context-learning-based (ICL-based) models have the potential towards the unification of medical lesion segmentation. However, due to their cross-fusion designs, existing ICL-based unified segmentation models fail to accurately localize lesions with low-matched reference sets.…

Cited by 0SourcePDFScholar
2024

Multi-view Aggregation Network for Dichotomous Image Segmentation

CVPR 2024highlight

Dichotomous Image Segmentation (DIS) has recently emerged towards high-precision object segmentation from high-resolution natural images. When designing an effective DIS model the main challenge is how to balance the semantic dispersion of high-resolution targets in the small receptive field and the…

2024

Spider: A Unified Framework for Context-dependent Concept Segmentation

ICML 2024poster

Different from the context-independent (CI) concepts such as human, car, and airplane, context-dependent (CD) concepts require higher visual understanding ability, such as camouflaged object and medical lesion. Despite the rapid advance of many CD understanding tasks in respective branches, the isol…

2023

Adaptive Illumination Mapping for Shadow Detection in Raw Images

ICCV 2023poster

Shadow detection methods rely on multi-scale contrast, especially global contrast, information to locate shadows correctly. However, we observe that the camera image signal processor (ISP) tends to preserve more local contrast information by sacrificing global contrast information during the raw-to-…

Cited by 15PDFcodeScholar
2023

Referring Image Segmentation Using Text Supervision

ICCV 2023poster

Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to localize the target object. Hence, we propose a novel weakly-…

Cited by 34PDFcodeScholar
2022

Self-Supervised Pretraining for RGB-D Salient Object Detection

AAAI 2022technical

Existing CNNs-Based RGB-D salient object detection (SOD) networks are all required to be pretrained on the ImageNet to learn the hierarchy features which helps provide a good initialization. However, the collection and annotation of large-scale datasets are time-consuming and expensive. In this pap…

2022

Zoom in and Out: A Mixed-Scale Triplet Network for Camouflaged Object Detection

CVPR 2022poster

The recently proposed camouflaged object detection (COD) attempts to segment objects that are visually blended into their surroundings, which is extremely complex and difficult in real-world scenarios. Apart from high intrinsic similarity between the camouflaged objects and their background, the obj…

Cited by 349PDFcodeScholar
2021

Encoder Fusion Network With Co-Attention Embedding for Referring Image Segmentation

CVPR 2021poster

Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature of each scale separately, which ignores the continuous guidan…

Cited by 196PDFScholar
2020

Bi-Directional Relationship Inferring Network for Referring Image Segmentation

CVPR 2020poster

Most existing methods do not explicitly formulate the mutual guidance between vision and language. In this work, we propose a bi-directional relationship inferring network (BRINet) to model the dependencies of cross-modal information. In detail, the vision-guided linguistic attention is used to lear…

Cited by 199PDFScholar
2020

Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection

ECCV 2020poster

The main purpose of RGB-D salient object detection (SOD) is how to better integrate and utilize cross-modal fusion information. In this paper, we explore these issues from a new perspective. We integrate the features of different modalities through densely connected structures and use their mixed fe…

2020

Suppress and Balance: A Simple Gated Network for Salient Object Detection

ECCV 2020poster

Most salient object detection approaches use U-Net or feature pyramid networks (FPN) as their basic structures. These methods ignore two key problems when the encoder exchanges information with the decoder: one is the lack of interference control between them, the other is without considering the di…

2019

Joint Learning of Saliency Detection and Weakly Supervised Semantic Segmentation

ICCV 2019poster

Existing weakly supervised semantic segmentation (WSSS) methods usually utilize the results of pre-trained saliency detection (SD) models without explicitly modelling the connections between the two tasks, which is not the most efficient configuration. Here we propose a unified multi-task learning f…

Cited by 246PDFcodeScholar
2019

Multi-Source Weak Supervision for Saliency Detection

CVPR 2019poster

The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency…

Cited by 227PDFcodeScholar
2018

Detect Globally, Refine Locally: A Novel Approach to Saliency Detection

CVPR 2018poster

Effective integration of contextual information is crucial for salient object detection. To achieve this, most existing methods based on 'skip' architecture mainly focus on how to integrate hierarchical features of Convolutional Neural Networks (CNNs). They simply apply concatenation or element-wise…

Cited by 507SourcePDFScholar
2017

A Stagewise Refinement Model for Detecting Salient Objects in Images

ICCV 2017poster

Deep convolutional neural networks (CNNs) have been successfully applied to a wide variety of problems in computer vision, including salient object detection. To detect and segment salient objects accurately, it is necessary to extract and combine high-level semantic features with low-level fine det…

Cited by 521PDFcodeScholar