← Search

Bingke Zhu

10 accepted papers

2026

AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection

AAAI 2026technical

Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized models, tailored for specific anomaly types like textural defects or logical errors, typically exhibit limited performanc

Cited by 0SourcePDFScholar
2026

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

AAAI 2026technical

Despite substantial progress in anomaly synthesis, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce

Cited by 0SourcePDFScholar
2026

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

CVPR 2026

Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, jointly optimizing these objectives in a single forward pass forces the contrastive branch to rely on randomly visible patches designed for reconstruct

Cited by 0SourceScholar
2025

FLARE: A Framework for Stellar Flare Forecasting Using Stellar Physical Properties and Historical Records

IJCAI 2025

Stellar flare events are critical observational samples for astronomical research; however, recorded flare events remain limited. Stellar flare forecasting can provide additional flare event samples to support research efforts. Despite this potential, no specialized models for stellar flare forecast

2025

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

ICASSP 2025accepted

Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal boundaries. Many methods ignore that different modalities often lack alignment, thereby introducing extra noise during modal…

Cited by 0SourceScholar
2025

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

ICCV 2025poster

The weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the limitations of the weakly-supervised and the deficiencies of the model architecture, existing methods are lacking in sim…

2025

UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection

CVPR 2025poster

Visual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical fields. Due to the domain gaps between these fields, existing VAD methods are typically tailored to each domain, with sp…

2024

AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models

AAAI 2024technical

Large Vision-Language Models (LVLMs) such as MiniGPT-4 and LLaVA have demonstrated the capability of understanding images and achieved remarkable performance in various visual tasks. Despite their strong abilities in recognizing common objects due to extensive training datasets, they lack specific d…