← Search

Jiaqing Zhang

3 accepted papers

2025

DiffCLIP: Few-shot Language-driven Multimodal Classifier

AAAI 2025technical

Visual language models like Contrastive Language-Image Pretraining (CLIP) have shown impressive performance in analyzing natural images with language information. However, these models often encounter challenges when applied to specialized domains such as remote sensing due to the limited availabili…

2024

E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection

NeurIPS 2024oral

Multimodal image fusion and object detection are crucial for autonomous driving. While current methods have advanced the fusion of texture details and semantic information, their complex training processes hinder broader applications. Addressing this challenge, we introduce E2E-MFD, a novel end-to-e…

2024

MDFL: Multi-Domain Diffusion-Driven Feature Learning

AAAI 2024technical

High-dimensional images, known for their rich semantic information, are widely applied in remote sensing and other fields. The spatial information in these images reflects the object's texture features, while the spectral information reveals the potential spectral representations across different ba…