← Search

Deng-Ping Fan

37 accepted papers

2026

AlignedNorm: Prompting Vision–Language Models via Coupled Prompt Field

ICML 2026poster

Prompt learning for vision-language models (VLMs) primarily follows end-to-end or decoupled routes to balance base and new task performance, but suffers a fundamental bottleneck: sample-wise optimization within task-specific feature spaces traps models in local optima, hindering global optimality. T…

Cited by 0SourceScholar
2025

AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial Patches

NeurIPS 2025poster

Cutting-edge works have demonstrated that text-to-image (T2I) diffusion models can generate adversarial patches that mislead state-of-the-art object detectors in the physical world, revealing detectors' vulnerabilities and risks. However, these methods neglect the T2I patches' attack effectiveness w…

Cited by 0SourcecodeScholar
2025

LawDIS: Language-Window-based Controllable Dichotomous Image Segmentation

ICCV 2025poster

We present LawDIS, a language-window-based controllable dichotomous image segmentation (DIS) framework that produces high-quality object masks. Our framework recasts DIS as an image-conditioned mask generation task within a latent diffusion model, enabling seamless integration of user controls. LawD…

2025

RUN: Reversible Unfolding Network for Concealed Object Segmentation

ICML 2025poster

Concealed object segmentation (COS) is a challenging problem that focuses on identifying objects that are visually blended into their background. Existing methods often employ reversible strategies to concentrate on uncertain regions but only focus on the mask level, overlooking the valuable of the…

2024

BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model

CVPR 2024poster

In this paper we address the challenge of image resolution variation for the Segment Anything Model (SAM). SAM known for its zero-shot generalizability exhibits a performance degradation when faced with datasets with varying image sizes. Previous approaches tend to resize the image to a fixed size o…

Cited by 20SourcePDFScholar
2024

LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented Diffusion

CVPR 2024poster

Camouflaged vision perception is an important vision task with numerous practical applications. Due to the expensive collection and labeling costs this community struggles with a major bottleneck that the species category of its datasets is limited to a small number of object species. However the ex…

2024

MaskFactory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation

NeurIPS 2024poster

Dichotomous Image Segmentation (DIS) tasks require highly precise annotations, and traditional dataset creation methods are labor intensive, costly, and require extensive domain expertise. Although using synthetic data for DIS is a promising solution to these challenges, current generative models an…

2024

Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes

CVPR 2024highlight

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes feature propagation or cross-frame attention to address these issues. By contrast we are the f…

2023

Indiscernible Object Counting in Underwater Scenes

CVPR 2023poster

Recently, indiscernible scene understanding has attracted a lot of attention in the vision community. We further advance the frontier of this field by systematically studying a new challenge named indiscernible object counting (IOC), the goal of which is to count objects that are blended with respec…

2023

Source-free Depth for Object Pop-out

ICCV 2023poster

Depth cues are known to be useful for visual perception. However, direct measurement of depth is often impracticable. Fortunately, though, modern learning-based methods offer promising depth maps by inference in the wild. In this work, we adapt such depth inference models for object segmentation usi…

Cited by 70PDFcodeScholar
2022

Highly Accurate Dichotomous Image Segmentation

ECCV 2022poster

"We present a systematic study on a new task called dichotomous image segmentation (DIS), which aims to segment highly accurate objects from natural images. To this end, we collected the first large-scale DIS dataset, called DIS5K, which contains 5,470 high-resolution (e.g., 2K, 4K or larger) images…

2022

Implicit Motion Handling for Video Camouflaged Object Detection

CVPR 2022poster

We propose a new video camouflaged object detection (VCOD) framework that can exploit both short-term dynamics and long-term temporal consistency to detect camouflaged objects from video frames. An essential property of camouflaged objects is that they usually exhibit patterns similar to the backgro…

Cited by 104PDFcodeScholar
2022

OSFormer: One-Stage Camouflaged Instance Segmentation with Transformers

ECCV 2022poster

"We present OSFormer, the first one-stage transformer framework for camouflaged instance segmentation (CIS). OSFormer is based on two key designs. First, we design a location-sensing transformer (LST) to obtain the location label and instance-aware parameters by introducing the location-guided queri…

2021

Camouflaged Object Segmentation With Distraction Mining

CVPR 2021poster

Camouflaged object segmentation (COS) aims to identify objects that are "perfectly" assimilate into their surroundings, which has a wide range of valuable applications. The key challenge of COS is that there exist high intrinsic similarities between the candidate objects and noise background. In thi…

Cited by 495PDFcodeScholar
2021

From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection Approach

CVPR 2021poster

Thanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its i…

Cited by 47PDFcodeScholar
2021

Full-Duplex Strategy for Video Object Segmentation

ICCV 2021poster

Appearance and motion are two important sources of information in video object segmentation (VOS). Previous methods mainly focus on using simplex solutions, lowering the upper bound of feature collaboration among and across these two cues. In this paper, we study a novel framework, termed the FSNet…

Cited by 180PDFcodeScholar
2021

Group Collaborative Learning for Co-Salient Object Detection

CVPR 2021poster

We present a novel group collaborative learning framework (GCNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations at group level based on the two necessary criteria: 1) intra-group compactness to better formulate the consistency among c…

Cited by 115PDFcodeScholar
2021

Kaleido-BERT: Vision-Language Pre-Training on Fashion Domain

CVPR 2021poster

We present a new vision-language (VL) pre-training model dubbed Kaleido-BERT, which introduces a novel kaleido strategy for fashion cross-modality representations from transformers. In contrast to random masking strategy of recent VL models, we design alignment guided masking to jointly focus more o…

Cited by 154PDFcodeScholar
2021

Mutual Graph Learning for Camouflaged Object Detection

CVPR 2021poster

Automatically detecting/segmenting object(s) that blend in with their surroundings is difficult for current models. A major challenge is that the intrinsic similarities between such foreground objects and background surroundings make the features extracted by deep model indistinguishable. To overcom…

Cited by 338PDFcodeScholar
2021

Probabilistic Model Distillation for Semantic Correspondence

CVPR 2021poster

Semantic correspondence is a fundamental problem in computer vision, which aims at establishing dense correspondences across images depicting different instances under the same category. This task is challenging due to large intra-class variations and a severe lack of ground truth. A popular solutio…

Cited by 26PDFcodeScholar
2021

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolutions

ICCV 2021poster

Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network useful for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification s…

Cited by 5162PDFcodeScholar
2021

RGB-D Saliency Detection via Cascaded Mutual Information Minimization

ICCV 2021poster

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to explicitly model the multi-modal information between RGB im…

Cited by 139PDFcodeScholar
2021

Simultaneously Localize, Segment and Rank the Camouflaged Objects

CVPR 2021poster

Camouflage is a key defence mechanism across species that is critical to survival. Common camouflage include background matching, imitating the color and pattern of the environment, and disruptive coloration, disguising body outlines. Camouflaged object detection (COD) aims to segment camouflaged ob…

Cited by 472PDFcodeScholar
2021

Specificity-Preserving RGB-D Saliency Detection

ICCV 2021poster

RGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preser…

Cited by 256PDFcodeScholar
2021

Uncertainty-Guided Transformer Reasoning for Camouflaged Object Detection

ICCV 2021poster

Spotting objects that are visually adapted to their surroundings is challenging for both humans and AI. Conventional generic / salient object detection techniques are suboptimal for this task because they tend to only discover easy and clear objects, while overlooking the difficult-to-detect ones wi…

Cited by 296PDFcodeScholar
2020

BBS-Net: RGB-D Salient Object Detection with a Bifurcated Backbone Strategy Network

ECCV 2020poster

Multi-level feature fusion is a fundamental topic in computer vision for detecting, segmenting, and classifying objects at various scales. When multi-level features meet multi-modal cues, the optimal fusion problem becomes a hot potato. In this paper, we make the first attempt to leverage the inhere…

2020

JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object Detection

CVPR 2020poster

This paper proposes a novel joint learning and densely-cooperative fusion (JL-DCF) architecture for RGB-D salient object detection. Existing models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constra…

Cited by 389PDFcodeScholar
2020

Taking a Deeper Look at Co-Salient Object Detection

CVPR 2020poster

Co-salient object detection (CoSOD) is a newly emerging and rapidly growing branch of salient object detection (SOD), which aims to detect the co-occurring salient objects in multiple images. However, existing CoSOD datasets often have a serious data bias, which assumes that each group of images con…

Cited by 100PDFScholar
2020

UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders

CVPR 2020oral

In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following…

Cited by 420PDFScholar
2019

Contrast Prior and Fluid Pyramid Integration for RGBD Salient Object Detection

CVPR 2019poster

The large availability of depth sensors provides valuable complementary information for salient object detection (SOD) in RGBD images. However, due to the inherent difference between RGB and depth information, extracting features from the depth channel using ImageNet pre-trained backbone models and…

Cited by 451PDFScholar
2019

EGNet: Edge Guidance Network for Salient Object Detection

ICCV 2019poster

Fully convolutional neural networks (FCNs) have shown their advantages in the salient object detection task. However, most existing FCNs-based methods still suffer from coarse object boundaries. In this paper, to solve this problem, we focus on the complementarity between salient edge information an…

Cited by 1302PDFScholar
2019

Multi-Level Context Ultra-Aggregation for Stereo Matching

CVPR 2019poster

Exploiting multi-level context information to cost volume can improve the performance of learning-based stereo matching methods. In recent years, 3-D Convolution Neural Networks (3-D CNNs) show the advantages in regularizing cost volume but are limited by unary features learning in matching cost com…

Cited by 137PDFScholar
2019

Scoot: A Perceptual Metric for Facial Sketches

ICCV 2019poster

While it is trivial for humans to quickly assess the perceptual similarity between two images, the underlying mechanism are thought to be quite complex. Despite this, the most widely adopted perceptual metrics today, such as SSIM and FSIM, are simple, shallow functions, and fail to consider many fac…

Cited by 57PDFScholar
2019

Shifting More Attention to Video Salient Object Detection

CVPR 2019oral

The last decade has witnessed a growing interest in video salient object detection (VSOD). However, the research community long-term lacked a well-established VSOD dataset representative of real dynamic scenes with high-quality annotations. To address this issue, we elaborately collected a visual-at…

Cited by 561PDFcodeScholar
2018

Salient Objects in Clutter: Bringing Salient Object Detection to the Foreground

ECCV 2018poster

We provide a comprehensive evaluation of salient object detection (SOD) models. Our analysis identifies a serious design bias of existing SOD datasets which assumes that each image contains at least one clearly outstanding salient object in low clutter. The design bias has led to a saturated high pe…

Cited by 380SourcePDFScholar
2017

Structure-Measure: A New Way to Evaluate Foreground Maps

ICCV 2017spotlight

Foreground map evaluation is crucial for gauging the progress of object segmentation algorithms, in particular in the filed of salient object detection where the purpose is to accurately detect and segment the most salient object in a scene. Several widely-used measures such as Area Under the Curve…

Cited by 1925PDFcodeScholar