← Search

Yuesheng Zhu

21 accepted papers

2026

DSFedMed: Dual-Scale Federated Medical Image Segmentation via Mutual Distillation Between Foundation and Lightweight Models

AAAI 2026technical

Foundation Models (FMs) have demonstrated strong generalization across diverse vision tasks. However, their deployment in federated settings is hindered by high computational demands, substantial communication overhead, and significant inference costs. We propose DSFedMed, a dual-scale federated fra

Cited by 0SourcePDFScholar
2026

Feature-Aware One-Shot Federated Learning via Hierarchical Token Sequences

AAAI 2026technical

One-shot federated learning (OSFL) reduces the communication cost and privacy risks of iterative federated learning by constructing a global model with a single round of communication. However, most existing methods struggle to achieve robust performance on real-world domains such as medical imaging

Cited by 0SourcePDFScholar
2026

SPOTR: Spatio-temporal Pooling One-Token Reconstruction for Universal Physiological Signal Self-supervised Learning

IJCAI 2026

Physiological signals such as EEG, ECG, and PPG are widely used in clinical monitoring. Recent self-supervised learning (SSL) methods offer an attractive way to leverage unlabeled recordings, yet they still fall short in practice. In particular, current SSL methods struggle across heterogeneous data

Cited by 0Scholar
2025

Optical Model-Driven Sharpness Mapping for Autofocus in Small Depth-of-Field and Severe Defocus Scenarios

ICCV 2025poster

Autofocus (AF) is essential for imaging systems, particularly in industrial applications such as automated optical inspection (AOI), where achieving precise focus is critical. Conventional AF methods rely on peak-searching algorithms that require dense focal sampling, making them inefficient in smal…

Cited by 0SourcePDFScholar
2025

Robust Image Hashing Based on Contrastive Masked Autoencoder with Weak-Strong Augmentation Alignment

AAAI 2025technical

Recently, numerous robust image hashing schemes have been developed for content identification. However, many of these schemes face the challenges of maintaining discrimination while simultaneously resisting large-scale attacks. In this paper, we propose a robust image hashing scheme based on Contra…

2025

Spiking Transformer with Spatial-Temporal Spiking Self-Attention

ICASSP 2025accepted

Spiking Neural Networks are celebrated for energy efficiency and biological plausibility. Building on Spiking Self-Attention (SSA), Spiking Transformers are extensively studied due to their exceptional performance. However, SSA focuses solely on spatial dimension at each time step, overlooking the c…

Cited by 0SourceScholar
2025

SpikingPoint: Rethinking Point as Spike for Efficient 3D Point Cloud Analysis

ICASSP 2025accepted

Spiking Neural Networks (SNNs), due to their unique spike-based inference mechanism, offer low power consumption and biological plausibility. As a fundamental technology for various real-world applications, 3D point cloud analysis faces significant challenges related to high computational overhead a…

Cited by 0SourceScholar
2024

Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling

AAAI 2024technical

Human evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the whole dataset in practice. However, different selection subsets will lead to diff…

2024

Enhanced Unsupervised Domain Adaptation with Dual-Attention Between Classification and Domain Alignment

ICASSP 2024accepted

Unsupervised Domain Adaptation (UDA) deals with transferring knowledge from labeled source domains to unlabeled target domains. This addresses the challenge of different distributions across domains, commonly known as domain shift. Numerous methods attempt to align distributions across domains while…

Cited by 0SourceScholar
2024

Spiking Transformer with Experts Mixture

NeurIPS 2024poster

Spiking Neural Networks (SNNs) provide a sparse spike-driven mechanism which is believed to be critical for energy-efficient deep learning. Mixture-of-Experts (MoE), on the other side, aligns with the brain mechanism of distributed and sparse processing, resulting in an efficient way of enhancing m…

Cited by 1SourcePDFScholar
2023

DeepVecFont-v2: Exploiting Transformers To Synthesize Vector Fonts With Higher Quality

CVPR 2023poster

Vector font synthesis is a challenging and ongoing problem in the fields of Computer Vision and Computer Graphics. The recently-proposed DeepVecFont achieved state-of-the-art performance by exploiting information of both the image and sequence modalities of vector fonts. However, it has limited capa…

2023

Robust Steganography without Embedding Based on Secure Container Synthesis and Iterative Message Recovery

IJCAI 2023poster

Synthesis-based steganography without embedding (SWE) methods transform secret messages to container images synthesised by generative networks, which eliminates distortions of container images and thus can fundamentally resist typical steganalysis tools. However, existing methods suffer from weak me…

Cited by 2SourcePDFScholar
2023

Spikformer: When Spiking Neural Network Meets Transformer

ICLR 2023poster

We consider two biologically plausible structures, the Spiking Neural Network (SNN) and the self-attention mechanism. The former offers an energy-efficient and event-driven paradigm for deep learning, while the latter has the ability to capture feature dependencies, enabling Transformer to achieve g…

2021

SERN: Stance Extraction and Reasoning Network for Fake News Detection

ICASSP 2021accepted

Fake news brings us panic and misunderstanding against the truth, especially under some unusual circumstances, such as the outbreak of COVID-19. It’s crucial to detect fake news on social media early to avoid further propagation. Previous methods manually label the stances implied in post-reply pair…

Cited by 0SourceScholar
2021

Semantic-Aware Context Aggregation for Image Inpainting

ICASSP 2021accepted

Recent attention-based image inpainting methods have made inspiring progress by propagating distant contextual information into holes. However, they tend to generate blurry contents since the propagation process is always misled by preliminarily-recovered holes features which are not well-inferred.…

Cited by 0SourceScholar
2020

Arnet: Attention-Based Refinement Network for Few-Shot Semantic Segmentation

ICASSP 2020accepted

Semantic segmentation is a challenging task for computer vision which aims to classify the objects from the pixel level. Previous methods based on deep learning have made some progress but the labeling work is very time-consuming. Few-shot semantic segmentation can alleviate this problem. In this pa…

Cited by 0SourceScholar
2020

Unsupervised Person Re-Identification Using Multi-Branch Feature Compensation Network and Link-Based Cluster Dissimilarity Metric

ICASSP 2020accepted

Feature extraction and label estimation are critical in unsupervised person re-identification (re-ID). Most previous works focus on acquiring high-layer semantic features and reckon without the lower-layer details lost in the learning process, which causes the extracted features to be less comprehen…

Cited by 0SourceScholar
2016

A Hole Filling Approach Based on Background Reconstruction for View Synthesis in 3D Video

CVPR 2016poster

The depth image based rendering (DIBR) plays a key role in 3D video synthesis, by which other virtual views can be generated from a 2D video and its depth map. However, in the synthesis process, the background occluded by the foreground objects might be exposed in the new view, resulting in some hol…

Cited by 76PDFScholar