← Search

Yukun Qian

6 accepted papers

2026

D²PPO: Diffusion Policy Policy Optimization with Dispersive Loss

AAAI 2026technical

Diffusion policies excel at robotic manipulation by naturally modeling multimodal action distributions in high-dimensional spaces. Nevertheless, diffusion policies suffer from diffusion representation collapse: semantically similar observations are mapped to indistinguishable features, ultimately im

Cited by 0SourcePDFScholar
2025

A Novel Compressive Compound Word Encoding and Independent Word Attention for Symbolic Music Generation

ICASSP 2025accepted

Symbolic music generation involves using symbolic encoding to represent music pieces as token sequences and using neural sequence models to create music by generating sequences of tokens. Symbolic encodings are primarily categorized into two types: independent word encoding and compound word encodin…

Cited by 0SourceScholar
2025

CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification

ICASSP 2025accepted

Sleep apnea is a common sleep disorder that, if untreated, can lead to serious health issues. Snoring is a typical symptom of sleep apnea and can be utilized to develop a noncontact automatic detection method for sleep apnea severity classification (SASC). However, due to patient heterogeneity, the…

Cited by 0SourceScholar
2025

Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank Adaptation

ICASSP 2025accepted

In real-world scenarios, accent variations often reduce Automatic Speech Recognition (ASR) accuracy. Addressing this typically involves a multi-task ASR and Accent Recognition (ASR-AR) framework, but there is limited research on optimizing task-specific feature extraction and enhancing ASR with AR i…

Cited by 0SourceScholar
2023

Half-Temporal and Half-Frequency Attention U2Net for Speech Signal Improvement

ICASSP 2023accepted

During communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to…

Cited by 0SourceScholar
2022

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

ICASSP 2022accepted

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Becaus…

Cited by 0SourceScholar