← Search

Xinzhong Zhu

16 accepted papers

2026

ExtendAttack: Attacking Servers of LRMs via Extending Reasoning

AAAI 2026technical

Large Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to maliciously occupy the resources of the servers, leading to a crash, like the DDoS attack in cyber. To this end, we propose a n

Cited by 0SourcePDFScholar
2026

Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection

ICLR 2026poster

Linear modeling methods like Mamba have been merged as the effective backbone for the 3D object detection task. However, previous Mamba-based methods utilize the bidirectional encoding for the whole non-empty voxel sequence, which contains abundant useless background information in the scenes. Thoug…

Cited by 0SourcecodeScholar
2026

Point Cloud Quantization Through Multimodal Prompting for 3D Understanding

AAAI 2026technical

Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current prototype-based approaches relying on trainable vectors or clustered centr

Cited by 0SourcePDFScholar
2026

RAC-DMVC: Reliability-Aware Contrastive Deep Multi-View Clustering Under Multi-Source Noise

AAAI 2026technical

Multi-view clustering (MVC), which aims to separate the multi-view data into distinct clusters in an unsupervised manner, is a fundamental yet challenging task. To enhance its applicability in real-world scenarios, this paper addresses a more challenging task: MVC under multi-source noises, includin

Cited by 0SourcePDFScholar
2025

Mamba YOLO: A Simple Baseline for Object Detection with State Space Model

AAAI 2025technical

Driven by the rapid development of deep learning technology, the YOLO series has set a new benchmark for real-time object detectors. Additionally, transformer-based structures have emerged as the most powerful solution in the field, greatly extending the model's receptive field and achieving signifi…

2025

MambaInst: Lightweight State Space Model for Real-Time Instance Segmentation

ICASSP 2025accepted

In this paper, we propose a lightweight and efficient state-space model-based instance segmentation network named MambaInst, which extracts deep semantic features through a LightSSM Block consisting of gating mechanisms and residual connectivity to model long-distance spatial dependencies with linea…

Cited by 0SourceScholar
2025

One-step Incomplete Multi-view Clustering based on Bipartite Graph Learning

ICASSP 2025accepted

Although previous graph-based multi-view clustering algorithms have made remarkable progress, most of them still face the following two limitations: 1. Many existing methods rely on k-means for the discretization of spectral embeddings, which cannot directly learn graphs with discrete cluster struct…

Cited by 0SourceScholar
2025

RemDet: Rethinking Efficient Model Design for UAV Object Detection

AAAI 2025technical

Object detection in Unmanned Aerial Vehicle (UAV) images has emerged as a focal area of research, which presents two significant challenges: i) objects are typically small and dense within vast images; ii) computational resource constraints render most models unsuitable for real-time deployment. Cur…

2025

RestorMamba: An Enhanced Synergistic State Space Model for Image Restoration

ICASSP 2025accepted

In this paper, we introduce an image inpainting method based on the State Space Model (SSM), named Restoration Mamba (RestorMamba). This approach incorporates effi-cient long-range dependency modeling within the network, which is particularly suited for the complexities of high-texture and high-reso…

Cited by 0SourceScholar
2025

Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image Fusion

ICML 2025poster

Multi-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well…

2025

Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts

NeurIPS 2025poster

Single-source Domain Generalized Object Detection (SDGOD), as a cutting-edge research topic in computer vision, aims to enhance model generalization capability in unseen target domains through single-source domain training. Current mainstream approaches attempt to mitigate domain discrepancies via d…

Cited by 0SourceScholar
2024

Scalable Multiple Kernel Clustering: Learning Clustering Structure from Expectation

ICML 2024poster

In this paper, we derive an upper bound of the difference between a kernel matrix and its expectation under a mild assumption. Specifically, we assume that the true distribution of the training data is an unknown isotropic Gaussian distribution. When the kernel function is a Gaussian kernel, and the…

Cited by 3SourcePDFScholar
2022

Align then Fusion: Generalized Large-scale Multi-view Clustering with Anchor Matching Correspondences

NeurIPS 2022accept

Multi-view anchor graph clustering selects representative anchors to avoid full pair-wise similarities and therefore reduce the complexity of graph methods. Although widely applied in large-scale applications, existing approaches do not pay sufficient attention to establishing correct correspondence…

2022

Highly-Efficient Incomplete Large-Scale Multi-View Clustering With Consensus Bipartite Graph

CVPR 2022poster

Multi-view clustering has received increasing attention due to its effectiveness in fusing complementary information without manual annotations. Most previous methods hold the assumption that each instance appears in all views. However, it is not uncommon to see that some views may contain some miss…

Cited by 143PDFcodeScholar
2019

DeFusionNET: Defocus Blur Detection via Recurrently Fusing and Refining Multi-Scale Deep Features

CVPR 2019poster

Defocus blur detection aims to detect out-of-focus regions from an image. Although attracting more and more attention due to its widespread applications, defocus blur detection still confronts several challenges such as the interference of background clutter, sensitivity to scales and missing bounda…

Cited by 88PDFScholar