← Search

Ying Hu

20 accepted papers

2025

A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision

ICASSP 2025accepted

Extracting singing melody from polyphonic music is an important topic in the field of music information retrieval. In this paper, we propose a singing melody extraction network consisting of five stacked multi-scale feature time-frequency aggregation (MF-TFA) modules. In the same network, deeper lay…

Cited by 0SourceScholar
2025

An Automatic Cutting Plane Planning Method Based on Multi-Objective Optimization for Robot-Assisted Laminectomy Surgery

RA-L 2025

Laminectomy represents an effective surgical procedure for the treatment of lumbar spinal stenosis. Due to the intricate anatomical structure of the lumbar spine, meticulous surgical path planning is essential to ensure the safety of the procedure and enhance the likelihood of successful outcomes. T

Cited by 6SourceScholar
2025

GDiffRetro: Retrosynthesis Prediction with Dual Graph Enhanced Molecular Representation and Diffusion Generation

AAAI 2025technical

Retrosynthesis prediction focuses on identifying reactants capable of synthesizing a target product. Typically, the retrosynthesis prediction involves two phases: Reaction Center Identification and Reactant Generation. However, we argue that most existing methods suffer from two limitations in the t…

2025

HANet: A Harmonic Attention-Based Network for Singing Melody Extraction from Polyphonic Music

ICASSP 2025accepted

Singing melody extraction from polyphonic music is a complex but important task in music information retrieval. Harmonic relationships have been shown to be crucial in this task, but most existing models based on Convolutional Neural Networks (CNNs) struggle to capture long-range harmonic dependenci…

Cited by 0SourceScholar
2024

A Decision-Making Algorithm for Robotic Breast Ultrasound High-Quality Imaging via Broad Reinforcement Learning From Demonstration

RA-L 2024

Robotic breast ultrasound (RBUS) aims to standardize breast ultrasonography, reduce the workload of sonographers, and provide high-quality ultrasound (US) images for subsequent diagnosis. In the process of RBUS screening, adjusting the US probe correctly and efficiently to acquire high-quality US im

Cited by 11SourceScholar
2024

Force-Position Hybrid Control for Robot Assisted Thoracic-Abdominal Puncture With Respiratory Movement

RA-L 2024

Percutaneous puncture is a widely used procedure in the diagnosis and therapy of cancer such as biopsy and ablation operations, while the organs in the thoracic and abdominal cavities are significantly affected by patients' respiratory movement. In this study, a robotic puncture system with respirat

Cited by 7SourceScholar
2024

Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker Verification

ICASSP 2024accepted

Incorporating frame-level phonetic information during the extraction of speaker embeddings has been shown to enhance the performance of speaker verification systems. However, previous studies have primarily relied on phonetic information obtained from pre-trained models of monolingual automatic spee…

Cited by 0SourceScholar
2024

Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models Function

NeurIPS 2024poster

Text-to-image diffusion models particularly Stable Diffusion, have revolutionized the field of computer vision. However, the synthesis quality often deteriorates when asked to generate images that faithfully represent complex prompts involving multiple attributes and objects. While previous studies…

2024

SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention

ICASSP 2024accepted

We propose a deep neural network with spectrogram matching and mutual attention (SMMA-Net) for audio clue-based target speaker extraction (TSE). To effectively use the auxiliary speech, we proposed spectrogram matching (SM) strategy and mutual attention (MA) block. We conducted all experiments on th…

Cited by 0SourceScholar
2024

SparseSSP: 3D Subcellular Structure Prediction from Sparse-View Transmitted Light Images

ECCV 2024oral

"Traditional fluorescence staining is phototoxic to live cells, slow, and expensive; thus, the subcellular structure prediction (SSP) from transmitted light (TL) images is emerging as a label-free, faster, low-cost alternative. However, existing approaches utilize 3D networks for one-to-one voxel le…

2024

Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE

NeurIPS 2024poster

Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical field remains a thorny challenge. To mitigate the tug-of-war problem of multi-m…

2023

Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation

ICASSP 2023accepted

Existing speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to miti…

Cited by 0SourceScholar
2022

Medical Ultrasound Image Quality Assessment for Autonomous Robotic Screening

RA-L 2022

Autonomous ultrasound scanning robots have attracted the attention of researchers, and the real-time quality assessment of ultrasound images is the key technology of them. Existing robot systems usually use pixel-level feature statistical methods such as grayscale, confidence map, etc. However, in c

Cited by 18SourceScholar
2022

Mining Hard Samples Locally And Globally For Improved Speech Separation

ICASSP 2022accepted

Speech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently goo…

Cited by 0SourceScholar
2021

Automatic Surgical Field of View Control in Robot-Assisted Nasal Surgery

RA-L 2021

In endoscopic nasal surgery, robots, rather than surgical assistants, can be introduced to hold endoscopes and act as the surgeon's third hand, which helps to reduce their operation burden. To address the problem of robot-assisted surgical field of view (FOV) acquisition in endoscopic nasal surgery,

Cited by 17SourceScholar
2021

Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification

ICASSP 2021accepted

We pursue an interpretable pitch tracking model and a jointly trained tone model for Mandarin tone classification. For pitch tracking, present deep learning based pitch model structure seldom considers the Viterbi decoding commonly implemented in prevalent manually designed pitch tracking algorithms…

Cited by 0SourceScholar
2021

Hybrid Adaptive Control Strategy for Continuum Surgical Robot Under External Load

RA-L 2021

Natural orifice transluminal endoscopic surgery (NOTES) has received significant attentions due to its minimal incision trauma compared with traditional multi-port robot assisted surgery. Continuum robot can be used in NOTES due to its high flexibility which can adapt to circuitous paths. However, t

Cited by 45SourceScholar
2017

A model of vertebral motion and key point recognition of drilling with force in robot-assisted spinal surgery

IROS 2017poster

Pedicle drilling is a crucial and high-risk process in spinal surgery. Due to the respiration and cardiac cycle, the position of spine would fluctuate during operations, which result in an increase of the difficulty in state recognition of pedicle drilling. To guarantee the safety and validity, a mo…

Cited by 8SourceScholar
2015

A novel optical tracking based tele-control system for tabletop object manipulation tasks

IROS 2015poster

For a robot serving in a complex environment such as in a restaurant, it is difficult to perform a task like tabletop object manipulation completely by itself, in that some information may be missing. An approach to deal with this is to use a tele-control system and method to control the robot or de…

Cited by 15SourceScholar