← Search

Hanyi Zhang

12 accepted papers

2025

Hybrid Deep Reinforcement Learning for Radio Tracer Localisation in Robotic-Assisted Radioguided Surgery

ICRA 2025

Radioguided surgery, such as sentinel lymph node biopsy, relies on the precise localization of radioactive targets by non-imaging gamma/beta detectors. Manual radioactive target detection based on visual display or audible indication of gamma level is highly dependent on the ability of the surgeon t

Cited by 1SourceScholar
2025

Multi-Modal and Multi-Attribute Generation of Single Cells with CFGen

ICLR 2025poster

Generative modeling of single-cell RNA-seq data is crucial for tasks like trajectory inference, batch effect removal, and simulation of realistic cellular data. However, recent deep generative models simulating synthetic single cells from noise operate on pre-processed continuous gene expression app…

2025

PiCNet: Physics-infused Convolution Network for Radar-Based Precipitation Nowcasting

ICASSP 2025accepted

Meteorological disasters, especially extreme precipitation, cause significant socioeconomic damage, highlighting the need for effective quantitative precipitation nowcasting. Existing methods, often data-driven and resource-intensive, struggle to capture the underlying physical laws of meteorology.…

Cited by 0SourceScholar
2024

Fast Context-Based Low-Light Image Enhancement via Neural Implicit Representations

ECCV 2024poster

"Current deep learning-based low-light image enhancement methods often struggle with high-resolution images, and fail to meet the practical demands of visual perception across diverse and unseen scenarios. In this paper, we introduce a novel approach termed CoLIE, which redefines the enhancement pro…

2024

Hybrid Robot for Percutaneous Needle Intervention Procedures: Mechanism Design and Experiment Verification

ICRA 2024poster

This paper presents a 6-DOF hybrid robot for percutaneous needle intervention procedures. The new robot combines the advantages of both serial robots and parallel robots, featuring compactness, high accuracy, and small footprint while overcoming the problems of the high cost of serial robots and the…

Cited by 0SourceScholar
2024

Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text Translation

ICASSP 2024accepted

End-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cros…

Cited by 0SourceScholar
2023

Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification

ICASSP 2023accepted

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is mo…

Cited by 0SourceScholar
2023

Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection

ICASSP 2023accepted

Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that levera…

Cited by 0SourceScholar
2023

Noise-Disentanglement Metric Learning for Robust Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglemen…

Cited by 0SourceScholar
2023

Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning

ICASSP 2023accepted

In self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory sig…

Cited by 0SourceScholar
2022

Learning Domain-Invariant Transformation for Speaker Verification

ICASSP 2022accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-le…

Cited by 0SourceScholar
2021

Meta-Learning for Cross-Channel Speaker Verification

ICASSP 2021accepted

Automatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channel…

Cited by 0SourceScholar