← Search

Wei-Qiang Zhang

20 accepted papers

2026

Backjump-on-Graph: Empowering LLMs with Reinforced Retrospective Exploration for Agentic KG Reasoning

ICML 2026poster

Grounding Large Language Models (LLMs) in Knowledge Graphs (KGs) has shown significant promise for complex Question Answering (QA) tasks. Since LLMs' limited context window cannot accommodate the sheer volume of large-scale KGs, existing work usually utilizes agents to reason on real-world KGs, whic…

Cited by 0SourceScholar
2026

Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

ICLR 2026poster

The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrain…

Cited by 0SourceScholar
2025

Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known Attributes

ICASSP 2025accepted

Adapting pre-trained models has become the dominant approach for anomalous sound detection (ASD), where classifying the attributes of machine working status is commonly chosen as the deputy task for fine-tuning. However, attributes might be intractable to collect for some machines, causing the label…

Cited by 0SourceScholar
2025

Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning

ICASSP 2025accepted

The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability o…

Cited by 0SourceScholar
2025

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

ACL 2025long

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual s…

2025

Improving Acoustic Scene Classification in Low-Resource Conditions

ICASSP 2025accepted

Acoustic Scene Classification (ASC) identifies an environment based on an audio signal. This paper explores ASC in low-resource conditions and proposes a novel model, DS-FlexiNet, which combines depthwise separable convolutions from MobileNetV2 with ResNet-inspired residual connections for a balance…

Cited by 0SourceScholar
2025

Integrating Pause Information with Word Embeddings in Language Models for Alzheimer's Disease Detection from Spontaneous Speech

ICASSP 2025accepted

Alzheimer’s disease (AD) is a progressive neurodegenerative disorder characterized by cognitive decline and memory loss. Early detection of AD is crucial for effective intervention and treatment. In this paper, we propose a novel approach to AD detection from spontaneous speech, which incorporates p…

Cited by 0SourceScholar
2025

Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning

ICASSP 2025accepted

Speech Emotion Recognition (SER) involves analyzing vocal expressions to determine the emotional state of speakers, where the comprehensive and thorough utilization of audio information is paramount. Therefore, we propose a novel approach on self-supervised learning (SSL) models that employs all ava…

Cited by 0SourceScholar
2025

WinStega: An Adaptive Robust Enhancement Framework for Generative Linguistic Steganography

ICASSP 2025accepted

With the increasing prevalence of surveillance, safeguarding personal privacy has become a critical concern. To protect privacy, various linguistic steganography methods have been developed to conceal private information within seemingly innocuous text for covert communication. However, these method…

Cited by 0SourceScholar
2024

Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection

ICASSP 2024accepted

Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages…

Cited by 0SourceScholar
2024

Whisper-Based Transfer Learning for Alzheimer Disease Classification: Leveraging Speech Segments with Full Transcripts as Prompts

ICASSP 2024accepted

Alzheimer’s disease (AD) is a neurodegenerative disorder that can lead to speech impairments. Early diagnosis is crucial for effective treatment, and speech-based diagnosis is currently a hot research topic. In this study, we explore the feasibility of transfer learning for Alzheimer’s disease detec…

Cited by 0SourceScholar
2023

Cross-Lingual Alzheimer's Disease Detection Based on Paralinguistic and Pre-Trained Features

ICASSP 2023accepted

We present our submission to the ICASSP-SPGC-2023 ADReSS-M Challenge Task, which aims to investigate which acoustic features can be generalized and transferred across languages for Alzheimer’s Disease (AD) prediction. The challenge consists of two tasks: one is to classify the speech of AD patients…

Cited by 0SourceScholar
2023

Unsupervised Anomaly Detection and Localization of Machine Audio: A Gan-Based Approach

ICASSP 2023accepted

Automatic detection of machine anomaly remains challenging for machine learning. We believe the capability of generative adversarial network (GAN) suits the need of machine audio anomaly detection, yet rarely has this been investigated by previous work. In this paper, we propose AEGAN-AD, a totally…

Cited by 0SourceScholar
2021

DeepRapper: Neural Rap Generation with Rhyme and Rhythm Modeling

ACL 2021long

Rap generation, which aims to produce lyrics and corresponding singing beats, needs to model both rhymes and rhythms. Previous works for rap generation focused on rhyming lyrics, but ignored rhythmic beats, which are important for rap performance. In this paper, we develop DeepRapper, a Transformer-…

2020

Dynamic Temporal Residual Learning for Speech Recognition

ICASSP 2020accepted

Long short-term memory (LSTM) networks have been widely used in automatic speech recognition (ASR). This paper proposes a novel dynamic temporal residual learning mechanism for LSTM networks to better explore temporal dependencies in sequential data. The temporal residual learning mechanism is imple…

Cited by 0SourceScholar
2020

Staged Training Strategy and Multi-Activation for Audio Tagging with Noisy and Sparse Multi-Label Data

ICASSP 2020accepted

Audio tagging aims to predict whether certain acoustic events occur in the audio clips. Due to the difficulty and huge cost of obtaining manually labeled data with high confidence, researchers begin to focus on audio tagging using a small set of manually-labeled data, and a larger set of noisy-label…

Cited by 0SourceScholar
2017

An LSTM-CTC based verification system for proxy-word based OOV keyword search

ICASSP 2017accepted

Proxy-word based out of vocabulary (OOV) keyword search has been proven to be quite effective in keyword search. In proxy-word based OOV keyword search, each OOV keyword is assigned several proxies and detections of the proxies are regarded as detections of the OOV keywords. However, the confidence…

Cited by 0SourceScholar
2017

Deep neural networks based speaker modeling at different levels of phonetic granularity

ICASSP 2017accepted

Recently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impa…

Cited by 0SourceScholar
2015

Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognition

ICASSP 2015accepted

Exploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold…

Cited by 0SourceScholar
2015

The THUEE system for the openKWS14 keyword search evaluation

ICASSP 2015accepted

The OpenKWS14 keyword search evaluation is one of the most challenging and influential evaluations in the field of speech recognition. Its goal is to build a high-performance keyword search system for a minority language with limited training data in a short period of time. We present the system of…

Cited by 0SourceScholar