← Search

Bob Zhang

20 accepted papers

2026

CondDiff-AMO: Integrating Conditional Diffusion Mechanism for Unified Amodal Mask Generation

AAAI 2026technical

Aiming to estimate the full extent of partially occluded objects, amodal segmentation is a critical capability for visual intelligence. Existing methods suffer from limitations in efficiency and precision, due to their reliance on auxiliary information or two-stage architectures. Furthermore, they l

Cited by 0SourcePDFScholar
2026

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

ICASSP 2026poster

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinforcement learning (RL) based post-training strategy for MLLMs in multi-image grou…

Cited by 0SourcePDFScholar
2026

Permutation-Consistent Variational Encoding for Incomplete Multi-View Multi-Label Classification

ICLR 2026poster

Incomplete multi-view multi-label learning is fundamentally an information integration problem under simultaneous view and label incompleteness. We introduce Permutation-Consistent Variational Encoding framework (PCVE) with an information bottleneck strategy, which learns variational representations…

Cited by 0SourceScholar
2026

Position: Reframing Hallucination: Latent Space Geodesics as a Pathway for Generative Discovery

ICML 2026poster

Current evaluation paradigms for generative models rely heavily on retrieval-based metrics such as exact match accuracy, creating a bottleneck particularly in domains requiring scientific discovery and creative reasoning. These metrics penalize any deviation from the training distribution, treating …

Cited by 0SourceScholar
2025

DCHM: Dynamic Collaboration of Heterogeneous Models Through Isomerism Learning in a Blockchain-Powered Federated Learning Framework

AAAI 2025technical

Solutions to time-varying problems are crucial for research areas such as predicting changes in human body shape over time. While recurrent neural networks have made significant advancements in this field, their reliance on centralized processing has led to challenges such as model silos and data is…

Cited by 0SourcePDFScholar
2025

Omni-Dimensional State Space Model-driven SAM for Pixel-level Anomaly Detection

IJCAI 2025

Pixel-level anomaly detection is indispensable in industrial defect detection and medical diagnosis. Recently, Segment Anything Model (SAM) has achieved promising results in many vision tasks. However, direct application of the SAM to pixel-level anomaly detection tasks results in unsatisfactory per

Cited by 0SourcePDFScholar
2025

Towards VLM-based Hybrid Explainable Prompt Enhancement for Zero-Shot Industrial Anomaly Detection

IJCAI 2025

Zero-Shot Industrial Anomaly Detection (ZSIAD) aims to identify and localize anomalies in industrial images from unseen categories. Owing to the powerful generalization capabilities, Vision-Language Models (VLMs) have achieved growing interest in ZSIAD. To guide the model toward understanding and lo

Cited by 0SourcePDFScholar
2025

Transforming Classification with Federated Learning on Blockchain: A Unique Model Integration Approach

ICASSP 2025accepted

The need for robust machine learning models is particularly evident in the realm of biological pattern recognition. Traditional centralized methods often struggle, as they frequently depend on large datasets that are challenging to gather due to stringent data privacy regulations. To address these l…

Cited by 0SourceScholar
2024

Open-Vocabulary Calibration for Fine-tuned CLIP

ICML 2024poster

Vision-language models (VLMs) have emerged as formidable tools, showing their strong capability in handling various open-vocabulary tasks in image recognition, text-driven visual content generation, and visual chatbots, to name a few. In recent years, considerable efforts and resources have been dev…

2023

Tensorized Incomplete Multi-View Clustering with Intrinsic Graph Completion

AAAI 2023technical

Most of the existing incomplete multi-view clustering (IMVC) methods focus on attaining a consensus representation from different views but ignore the important information hidden in the missing views and the latent intrinsic structures in each view. To tackle these issues, in this paper, a unified…

2022

Generative Adaptive Convolutions for Real-World Noisy Image Denoising

AAAI 2022technical

Recently, deep learning techniques are soaring and have shown dramatic improvements in real-world noisy image denoising. However, the statistics of real noise generally vary with different camera sensors and in-camera signal processing pipelines. This will induce problems of most deep denoisers for…

Cited by 31SourcePDFScholar
2021

Unified Tensor Framework for Incomplete Multi-view Clustering and Missing-view Inferring

AAAI 2021technical

In this paper, we propose a novel method, referred to as incomplete multi-view tensor spectral clustering with missing-view inferring (IMVTSC-MVI) to address the challenging multi-view clustering problem with missing views. Different from the existing methods which commonly focus on exploring the ce…

Cited by 157SourcePDFScholar
2020

A Noninvasive Method to Detect Diabetes Mellitus and Lung Cancer Using the Stacked Sparse autoencoder

ICASSP 2020accepted

Diabetes mellitus and lung cancer are two of the most common fatal diseases in the world, causing considerable deaths every year. However, it is not easy to detect diabetes mellitus and lung cancer efficiently--needing professional medical instruments such as a CT and a qualified individual to perfo…

Cited by 0SourceScholar
2020

CDIMC-net: Cognitive Deep Incomplete Multi-view Clustering Network

IJCAI 2020poster

In recent years, incomplete multi-view clustering, which studies the challenging multi-view clustering problem on missing views, has received growing research interests. Although a series of methods have been proposed to address this issue, the following problems still exist: 1) Almost all of the ex…

Cited by 0SourcePDFScholar
2020

Discriminant and Sparsity Based Least Squares Regression with l1 Regularization for Feature Representation

ICASSP 2020accepted

Least squares regression (LSR) has two main issues that greatly limits the improvement of performance: 1) The target matrix is too rigid leading to a large regression error; 2) the underlying geometric structure of the training data is often ignored to learn a more discriminative projection matrix.…

Cited by 0SourceScholar
2019

Learning Discriminative Finger-knuckle-print Descriptor

ICASSP 2019accepted

Direction information has been intensively investigated for Finger-Knuckle-Print (FKP) recognition. However, most existing direction-based KFP recognition methods are handcrafted, which are heuristic and require too much prior knowledge to engineer them. In this paper, we propose a discriminative di…

Cited by 0SourceScholar