ACD-CLIP: DECOUPLING REPRESENTATION AND DYNAMIC FUSION FOR ZERO-SHOT ANOMALY DETECTION
Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense prediction and employ inflexible feature fusion paradigms. We address these limitations through an Architectural Co-Des…