← Search

Jiqing Han

21 accepted papers

2025

InfoMin-based Query Embedding Optimization For Query-based Universal Sound Separation

ICASSP 2025accepted

The query-based universal sound separation (QUSS) has been addressed, aiming to perform the separation of specific sound sources based on a given query. Most of existed methods focus on the improvement of separation models, ignoring the influence of category-conditioned query embedding distribution…

Cited by 0SourceScholar
2025

Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets

EMNLP 2025

Language Model (LM)-based Text-to-Speech (TTS) systems often generate hallucinated speech that deviates from input text. Existing mitigation strategies either demand excessive training resources or introduce significant inference latency. In this paper, we propose GFlOwNet-guided distribution Alignm

2024

Contrastive Loss Based Frame-Wise Feature Disentanglement for Polyphonic Sound Event Detection

ICASSP 2024accepted

Overlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critical reason is that these methods represent overlapping events using shared and entangled frame-wise features, which degrad…

Cited by 0SourceScholar
2024

Modeling Quasi-Periodic Dependency via Self-Supervised Pre-Training for Respiratory Sound Classification

ICASSP 2024accepted

Despite the success of self-supervised respiratory sound classification methods, they do not consider that respiratory sounds are quasi-periodic signals with repetitive patterns in successive breaths, which is vital for distinguishing respiratory sounds from non-quasi-periodic sounds like noises. Th…

Cited by 0SourceScholar
2023

Graph-Based Spectro-Temporal Dependency Modeling for Anti-Spoofing

ICASSP 2023accepted

A great deal of recent research reveals that artifacts introduced by spoofing algorithms reside in specific frequency subbands or temporal segments. Therefore, the performance of spoofing detection can be improved by focusing on these regions. However, it is difficult for the detection system to cho…

Cited by 0SourceScholar
2023

Sentiment Knowledge Enhanced Self-supervised Learning for Multimodal Sentiment Analysis

ACL 2023findings

Multimodal Sentiment Analysis (MSA) has made great progress that benefits from extraordinary fusion scheme. However, there is a lack of labeled data, resulting in severe overfitting and poor generalization for supervised models applied in this field. In this paper, we propose Sentiment Knowledge Enh…

Cited by 11SourcePDFScholar
2023

Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection

ICASSP 2023accepted

Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency repres…

Cited by 0SourceScholar
2023

Using Auxiliary Tasks In Multimodal Fusion of Wav2vec 2.0 And Bert for Multimodal Emotion Recognition

ICASSP 2023accepted

The lack of data and the difficulty of multimodal fusion have always been challenges for multimodal emotion recognition (MER). In this paper, we propose to use pre-trained models as upstream network, wav2vec 2.0 for audio modality and BERT for text modality, and finetune them in downstream task of M…

Cited by 0SourceScholar
2022

Exploring Transformer's Potential on Automatic Piano Transcription

ICASSP 2022accepted

Most recent research about automatic music transcription (AMT) uses convolutional neural networks and recurrent neural networks to model the mapping from music signals to symbolic notation. Based on a high-resolution piano transcription system, we explore the possibility of incorporating another pow…

Cited by 0SourceScholar
2021

Capturing Temporal Dependencies Through Future Prediction for CNN-Based Audio Classifiers

ICASSP 2021accepted

This paper focuses on the problem of temporal dependency modeling in the CNN-based models for audio classification tasks. To capture audio temporal dependencies using CNNs, we take a different approach from the purely architecture-induced method and explicitly encode temporal dependencies into the C…

Cited by 0SourceScholar
2019

Furcax: End-to-end Monaural Speech Separation Based on Deep Gated (De)convolutional Neural Networks with Adversarial Example Training

ICASSP 2019accepted

Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain. Such an approach will result in limited perceptual score, s…

Cited by 0SourceScholar
2018

Deep Neural Network Based Discriminative Training for I-Vector/PLDA Speaker Verification

ICASSP 2018accepted

In the studies of i-vector based speaker verification, the discriminative training of probabilistic linear discriminative analysis (PLDA) model has been proven to be an effective way to improve performance. This paper focuses on using a deep neural network (DNN) to strengthen the original discrimina…

Cited by 0SourceScholar
2016

Realistic human action recognition: When deep learning meets VLAD

ICASSP 2016accepted

Human action recognition from realistic scenarios is extremely challenging due to large intra-class variation and complex background clutters. In this paper, by leveraging the strength of deep learning and vector of locally aggregated descriptors (VLAD), we propose a new methods for human action rec…

Cited by 0SourceScholar