← Search

Zhikang Zhang

9 accepted papers

2026

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

CVPR 2026

Large vision-language models (VLMs) show strong multimodal understanding but still struggle with 3D spatial reasoning, such as distance estimation, size comparison, and cross-view consistency. Existing 3D-aware methods either depend on auxiliary 3D information or enhance RGB-only VLMs with geometry

Cited by 0SourceScholar
2023

Automatic Error Detection in Integrated Circuits Image Segmentation: A Data-Driven Approach

ICASSP 2023accepted

Due to the complicated nanoscale structures of current integrated circuits(IC) builds and low error tolerance of IC image segmentation tasks, most existing automated IC image segmentation approaches require human experts for visual inspection to ensure correctness, which is one of the major bottlene…

Cited by 0SourceScholar
2023

Enhanced Low-Resolution LiDAR-Camera Calibration via Depth Interpolation and Supervised Contrastive Learning

ICASSP 2023accepted

Motivated by the increasing application of low-resolution LiDAR, we target the problem of low-resolution LiDAR-camera calibration in this work. The main challenges are two-fold: sparsity and noise in point clouds. To address the problem, we propose to apply depth interpolation to increase the point…

Cited by 0SourceScholar
2023

TransUPR: A Transformer-based Plug-and-Play Uncertain Point Refiner for LiDAR Point Cloud Semantic Segmentation

IROS 2023poster

Common image-based LiDAR point cloud semantic segmentation (LiDAR PCSS) approaches have bottlenecks resulting from the boundary-blurring problem of convolution neural networks (CNNs) and quantitation loss of spherical projection. In this work, we propose a transformer-based plug-and-play uncertain p…

Cited by 2SourceScholar
2022

An Experimental Study on Transferring Data-Driven Image Compressive Sensing to Bioelectric Signals

ICASSP 2022accepted

The emerging area of bioelectric signal compressive sensing(CS) has shown great potential in health care applications. However, improving the reconstruction accuracy of compressively sensed bioelectric signals remains a challenging problem. In recent years, data-driven image CS methods have achieved…

Cited by 0SourceScholar
2020

Cra: A Generic Compression Ratio Adapter for End-To-End Data-Driven Image Compressive Sensing Reconstruction Frameworks

ICASSP 2020accepted

End-to-end data-driven image compressive sensing reconstruction (EDCSR) frameworks achieve state-of-the-art reconstruction performance in terms of reconstruction speed and accuracy. However, due to their end-to-end nature, existing EDCSR frameworks can not adapt to a variable compression ratio (CR).…

Cited by 0SourceScholar
2018

LAPRAN: A Scalable Laplacian Pyramid Reconstructive Adversarial Network for Flexible Compressive Sensing Reconstruction

ECCV 2018poster

This paper addresses the single-image compressive sensing (CS) and reconstruction problem. We propose a scalable Laplacian pyramid reconstructive adversarial network (LAPRAN) that enables high-fidelity, flexible and fast CS images reconstruction. LAPRAN progressively reconstructs an image following…