ICASSP 2025accepted0 citations

Quad-Net: Melspectrogram Vocoder with Convolutional Layers Restricted by the Quadrature Mirror Filter for Perfect Reconstruction

Nam-Seok Song, Joon-Hyuk Chang

Abstract

Recently, neural vocoders have applied signal processing methods to synthesize speech to reduce computational complexity. However, most methods lack the benefits of a data-driven approach and the flexibility of hyper-parameters, such as filter length, because they rely on fixed signal processing filters. In this paper, we introduce Quad-Net, a network that includes restricted convolutional layers shaped by quadrature mirror synthesis filter banks. It is optimized with a perfect reconstruction loss derived from perfect reconstruction filter banks. This enables us to control filter lengths and degrees of data-drivenness. The results show that the filter parameters trained in our model exhibit characteristics similar to those of other signal processing methods with lower parameters. Furthermore, by increasing the filter length of Quad-Net, we can obtain filters that have complex frequency responses. It shows that a new approach enables the design of more complex filters that are adaptive to neural networks, diverging from previous methods.

BibTeX
@inproceedings{icassp2025_quadnetmelspectr,
  title = {Quad-Net: Melspectrogram Vocoder with Convolutional Layers Restricted by the Quadrature Mirror Filter for Perfect Reconstruction},
  author = {Nam-Seok Song and Joon-Hyuk Chang},
  booktitle = {ICASSP 2025},
  year = {2025}
}