← Search

Zhen Cao

6 accepted papers

2026

Denoise and Align: Towards Source-Free UDA for Robust Panoramic Semantic Segmentation

CVPR 2026

Panoramic semantic segmentation is pivotal for comprehensive 360deg scene understanding in critical applications like autonomous driving and virtual reality. However, progress in this domain is constrained by two key challenges: the severe geometric distortions inherent in panoramic projections and

Cited by 0SourcecodeScholar
2026

Narrowing the ANN–SNN Gap for 1D Signal Classification with Multi-Scale Temporal Encoding and Sparsity-Regularized Transform Encoding

ICML 2026poster

Spiking neural networks (SNNs) promise energy-efficient inference, yet on static vision benchmarks they often trail matched ANNs under short simulation horizons. Under a matched-backbone and matched-budget protocol without extra tricks, we find that this ANN-SNN accuracy gap is consistently smaller …

Cited by 0SourceScholar
2026

Optimization Method for Surrogate Function in Spiking Neural Networks Based on Membrane Potential Distribution

AAAI 2026technical

Spiking Neural Networks (SNNs) offer promising energy efficiency and temporal sparsity for edge intelligence, but their training remains difficult due to gradient mismatch, membrane potential drift, and discretization errors. In this paper, we propose a membrane potential-guided surrogate optimizati

Cited by 0SourcePDFScholar
2025

TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding

EMNLP 2025

Text-rich images are ubiquitous in real-world applications, serving as a critical medium for conveying complex information and facilitating accessibility.Despite recent advances driven by Multimodal Large Language Models (MLLMs), existing benchmarks suffer from limited scale, fragmented scenarios, a

Cited by 0SourcePDFScholar
2024

Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion

ECCV 2024poster

"∗ Equal contribution Corresponding authorIn this paper, we explore a novel framework, EGIInet (Explicitly Guided Information Interaction Network), a model for View-guided Point cloud Completion (ViPC) task, which aims to restore a complete point cloud from a partial one with a single view image. In…

2023

KT-Net: Knowledge Transfer for Unpaired 3D Shape Completion

AAAI 2023technical

Unpaired 3D object completion aims to predict a complete 3D shape from an incomplete input without knowing the correspondence between the complete and incomplete shapes. In this paper, we propose the novel KTNet to solve this task from the new perspective of knowledge transfer. KTNet elaborates a te…