← Search

Da Zhang

12 accepted papers

2026

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

AAAI 2026technical

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these

Cited by 0SourcePDFScholar
2026

Exploring the Underwater World Segmentation without Extra Training

CVPR 2026

Accurate segmentation of marine organisms is vital for biodiversity monitoring and ecological assessment, yet existing datasets and models remain largely limited to terrestrial scenes. To bridge this gap, we introduce **AquaOV255**, the first large-scale and fine-grained underwater segmentation data

Cited by 0SourcecodeScholar
2026

IntroSVG: Learning from Rendering Feedback for Text-to-SVG Generation via an Introspective Generator-Critic Framework

CVPR 2026

Scalable Vector Graphics (SVG) are central to digital design due to their inherent scalability and editability. Despite significant advancements in content generation enabled by Visual Language Models (VLMs), existing text-to-SVG generation methods are limited by a core challenge: the autoregressive

Cited by 0SourceScholar
2026

MedMamba: Multi-View State Space Models with Adaptive Graph Learning for Medical Time Series Classification

ICML 2026poster

Medical time series are central to healthcare, enabling continuous monitoring and supporting timely clinical decisions. Despite recent progress, existing methods struggle to jointly model local-global dynamics and handle nonstationarities like baseline drift, while often failing to capture latent ch…

Cited by 0SourceScholar
2025

ADD: A Detection Method for Image-Processing Adversarial Defenses

ICASSP 2025accepted

Many studies have demonstrated the vulnerability of Deep Neural Networks (DNNs) to adversarial attacks. While numerous research efforts have proposed high-performance adversarial attacks and defenses, there is a lack of research regarding the detection of defenses used by models. We have observed th…

Cited by 0SourceScholar
2022

INT: Towards Infinite-Frames 3D Detection with an Efficient Framework

ECCV 2022poster

"It is natural to construct a multi-frame instead of a single-frame 3D detector for a continuous-time stream. Although increasing the number of frames might improve performance, previous multi-frame studies only used very limited frames to build their systems due to the dramatically increased comput…

2022

LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object Detection

CVPR 2022poster

LiDAR and camera are two common sensors to collect data in time for 3D object detection under the autonomous driving context. Though the complementary information across sensors and time has great potential of benefiting 3D perception, taking full advantage of sequential cross-sensor data still rema…

Cited by 37PDFScholar
2019

MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment

CVPR 2019poster

This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex temporal dependencies, which often happens in real scenarios. We identify two cr…

Cited by 372PDFScholar
2017

Multimodal Transfer: A Hierarchical Deep Convolutional Neural Network for Fast Artistic Style Transfer

CVPR 2017poster

Transferring artistic styles onto everyday photographs has become an extremely popular task in both academia and industry. Recently, offline training has replaced online iterative optimization, enabling nearly real-time stylization. When those stylization networks are applied directly to high-reso…

Cited by 217PDFScholar