← Search

Wenfu Wang

3 accepted papers

2026

Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning

AAAI 2026technical

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded su

Cited by 0SourcePDFScholar
2017

Combining unidirectional long short-term memory with convolutional output layer for high-performance speech synthesis

ICASSP 2017accepted

In this paper, we target improving the accuracy of acoustic modelling for statistical parametric speech synthesis (SPSS) and introduce the convolutional neural network (CNN) due to its powerful capacity in locality modelling. A novel model architecture combining unidirectional long short-term memory…

Cited by 0SourceScholar
2016

Gating recurrent mixture density networks for acoustic modeling in statistical parametric speech synthesis

ICASSP 2016accepted

Though recurrent neural networks (RNNs) using long short-term memory (LSTM) units can address the issue of long-span dependencies across the linguistic inputs and have achieved the state-of-the-art performance for statistical parametric speech synthesis (SPSS), another limitation of the intrinsic un…

Cited by 0SourceScholar