← Search

Bingfeng Zhang

15 accepted papers

2026

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

AAAI 2026technical

3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. Howe

Cited by 0SourcePDFScholar
2026

TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection

CVPR 2026

Co-salient Object Detection (CoSOD) aims to segment salient objects that consistently appear across a group of related images. Despite the notable progress achieved by recent training-based approaches, they still remain constrained by the closed-set datasets and exhibit limited generalization. Howev

Cited by 0SourcecodeScholar
2026

The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVA

CVPR 2026

Multimodal Large Language Models (MLLMs) like LLaVA have demonstrated remarkable capabilities in multi-modal understanding and generation. This success motivates us to investigate whether the inherent prior knowledge embedded within such MLLMs contains sufficient spatial awareness for dense predicti

Cited by 0SourcecodeScholar
2025

A Training-free Synthetic Data Selection Method for Semantic Segmentation

AAAI 2025technical

Training semantic segmenter with synthetic data has been attracting great attention due to its easy accessibility and huge quantities. Most previous methods focused on producing large-scale synthetic image-annotation samples and then training the segmenter with all of them. However, such a solution…

2025

Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Training-free open-vocabulary semantic segmentation has advanced with vision-language models like CLIP, which exhibit strong zero-shot abilities. However, CLIP's attention mechanism often wrongly emphasises specific image tokens, namely outliers, which results in irrelevant over-activation. Existing…

2024

Adaptive Bidirectional Displacement for Semi-Supervised Medical Image Segmentation

CVPR 2024poster

Consistency learning is a central strategy to tackle unlabeled data in semi-supervised medical image segmentation (SSMIS) which enforces the model to produce consistent predictions under the perturbation. However most current approaches solely focus on utilizing a specific single perturbation which…

2024

Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation

CVPR 2024highlight

Weakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model while there is no attempt to apply the CLIP model as the backbone to directly segment…

2024

PSDPM: Prototype-based Secondary Discriminative Pixels Mining for Weakly Supervised Semantic Segmentation

CVPR 2024poster

Image-level Weakly Supervised Semantic Segmentation (WSSS) has received increasing attention due to its low annotation cost. Class Activation Mapping (CAM) generated through classifier weights in WSSS inevitably ignores certain useful cues while the CAM generated through class prototypes can allevia…

2024

Rethinking Prior Information Generation with CLIP for Few-Shot Segmentation

CVPR 2024poster

Few-shot segmentation remains challenging due to the limitations of its labeling information for unseen classes. Most previous approaches rely on extracting high-level feature maps from the frozen visual encoder to compute the pixel-wise similarity as a key prior guidance for the decoder. However su…

2023

Hunting Sparsity: Density-Guided Contrastive Learning for Semi-Supervised Semantic Segmentation

CVPR 2023poster

Recent semi-supervised semantic segmentation methods combine pseudo labeling and consistency regularization to enhance model generalization from perturbation-invariant training. In this work, we argue that adequate supervision can be extracted directly from the geometry of feature space. Inspired by…

2022

CARD: Semi-supervised Semantic Segmentation via Class-agnostic Relation based Denoising

IJCAI 2022poster

Recent semi-supervised semantic segmentation methods focus on mining extra supervision from unlabeled data by generating pseudo labels. However, noisy labels are inevitable in this process which prevent effective self-supervision. This paper proposes that noisy labels can be corrected based on seman…

Cited by 9SourcePDFScholar
2022

Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object Detection

CVPR 2022poster

Co-salient object detection, with the target of detecting co-existed salient objects among a group of images, is gaining popularity. Recent works use the attention mechanism or extra information to aggregate common co-salient features, leading to incomplete even incorrect responses for target object…

Cited by 76PDFcodeScholar
2021

Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency Coherence

AAAI 2021technical

Sparse labels have been attracting much attention in recent years. However, the performance gap between weakly supervised and fully supervised salient object detection methods is huge, and most previous weakly supervised works adopt complex training methods with many bells and whistles. In this work…

2020

Fast Template Matching and Update for Video Object Tracking and Segmentation

CVPR 2020poster

In this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the se…

Cited by 82PDFcodeScholar