← Search

Jen-Yu Liu

7 accepted papers

2026

AVEX: What Matters for Animal Vocalization Encoding

ICLR 2026poster

Bioacoustics, the study of sounds produced by living organisms, plays a vital role in conservation, biodiversity monitoring, and behavioral studies. Many tasks in this field, such as species, individual, and behavior classification and detection, are well-suited to machine learning. However, they of…

Cited by 0SourcecodeScholar
2025

Biodenoising: Animal Vocalization Denoising without Access to Clean Data

ICASSP 2025accepted

Animal vocalization denoising is a task similar to human speech enhancement, which is relatively well-studied. In contrast to the latter, it comprises a higher diversity of sound production mechanisms and recording environments, and this higher diversity is a challenge for existing models. Adding to…

Cited by 8SourceScholar
2023

BEANS: The Benchmark of Animal Sounds

ICASSP 2023accepted

The use of machine learning (ML) based techniques has become increasingly popular in the field of bioacoustics over the last years. Fundamental requirement for the successful application of ML based techniques are curated, agreed upon, high-quality datasets and benchmark tasks to be learned on a giv…

Cited by 72SourceScholar
2022

KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE Using Mel-Spectrograms

ICASSP 2022accepted

In this paper, we propose a novel neural network model called KaraSinger for a less-studied singing voice synthesis (SVS) task named score-free SVS, in which the prosody and melody are spontaneously decided by machine. KaraSinger comprises a vector-quantized variational autoencoder (VQ-VAE) that com…

Cited by 0SourceScholar
2021

Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs

AAAI 2021technical

To apply neural sequence models such as the Transformers to music generation tasks, one has to represent a piece of music by a sequence of tokens drawn from a finite set of pre-defined vocabulary. Such a vocabulary usually involves tokens of various types. For example, to describe a musical note, on…

2017

Revisiting the problem of audio-based hit song prediction using convolutional neural networks

ICASSP 2017accepted

Being able to predict whether a song can be a hit has important applications in the music industry. Although it is true that the popularity of a song can be greatly affected by external factors such as social and commercial influences, to which degree audio features computed from musical signals (wh…

Cited by 0SourceScholar
2017

Weakly-supervised audio event detection using event-specific Gaussian filters and fully convolutional networks

ICASSP 2017accepted

Audio event detection aims at discovering the elements inside an audio clip. In addition to labeling the clips with the audio events, we want to find out the temporal locations of these events. However, creating clearly annotated training data can be time-consuming. Therefore, we provide a model bas…

Cited by 0SourceScholar