← Search

Zhiyao Duan

29 accepted papers

2025

Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges

ICASSP 2025accepted

In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and A…

Cited by 1SourceScholar
2024

Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech

ICASSP 2024accepted

Dimensional representations of speech emotions such as the arousal-valence (AV) representation provide a continuous and fine-grained description and control than their categorical counterparts. They have wide applications in tasks such as dynamic emotion understanding and expressive text-to-speech s…

Cited by 0SourceScholar
2024

SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription

ICASSP 2024accepted

Guitar tablature is a form of music notation widely used among guitarists. It captures not only the musical content of a piece, but also its implementation and ornamentation on the instrument. Guitar Tablature Transcription (GTT) is an important task with broad applications in music education, compo…

Cited by 0SourceScholar
2022

A Novel 1D State Space for Efficient Music Rhythmic Analysis

ICASSP 2022accepted

Inferring music time structures has a broad range of applications in music production, processing and analysis. Scholars have proposed various methods to analyze different aspects of time structures, such as beat, downbeat, tempo and meter. Many state-of-the-art (SOFA) methods, however, are computat…

Cited by 0SourceScholar
2022

A Study of The Robustness of Raw Waveform Based Speaker Embeddings Under Mismatched Conditions

ICASSP 2022accepted

In this paper, we conduct a cross-dataset study on parametric and non-parametric raw-waveform based speaker embeddings through speaker verification experiments. In general, we observe a more significant performance degradation of these raw-waveform systems compared to spectral based systems. We then…

Cited by 0SourceScholar
2022

Progressive Teacher-Student Training Framework for Music Tagging

ICASSP 2022accepted

Music tagging is the task of predicting multiple tags of a music excerpt, and plays an important role in modern music recommendation systems. To obtain superior performance, recent approaches of music tagging focus on developing sophisticated models or exploiting additional multi-modal information.…

Cited by 0SourceScholar
2021

Don't Look Back: An Online Beat Tracking Method Using RNN and Enhanced Particle Filtering

ICASSP 2021accepted

Online beat tracking (OBT) has always been a challenging task. Due to the inaccessibility of future data and the need to make inference in real-time. We propose Don’t Look back! (DLB), a novel approach optimized for efficiency when performing OBT. DLB feeds the activations of a unidirectional RNN in…

Cited by 0SourceScholar
2021

Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFs

NeurIPS 2021poster

Piano transcription systems are typically optimized to estimate pitch activity at each frame of audio. They are often followed by carefully designed heuristics and post-processing algorithms to estimate note events from the frame-level predictions. Recent methods have also framed piano transcription…

2020

End-To-End Generation of Talking Faces from Noisy Speech

ICASSP 2020accepted

Acoustic cues are not the only component in speech communication; if the visual counterpart is present, it is shown to benefit speech comprehension. In this work, we propose an end-to-end (no pre- or post-processing) system that can generate talking faces from arbitrarily long noisy speech. We propo…

Cited by 0SourceScholar
2019

Hierarchical Cross-Modal Talking Face Generation With Dynamic Pixel-Wise Loss

CVPR 2019poster

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we propose first to transfer audio to high-level structure, i.e…

Cited by 490PDFcodeScholar
2018

Audio-Visual Event Localization in Unconstrained Videos

ECCV 2018poster

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio-Visual Event (AVE) dataset to systemically investigate three temporal localizati…

Cited by 575SourcePDFScholar
2018

Joint Speaker Diarization and Recognition Using Convolutional and Recurrent Neural Networks

ICASSP 2018accepted

Speaker diarization (detecting who-spoke-when using relative identity labels) and speaker recognition (detecting absolute identity labels without timing) are different but related tasks that often need to be completed simultaneously in many scenarios. Traditional methods, however, address them indep…

Cited by 0SourceScholar
2018

Unsupervised Learning Approach to Feature Analysis for Automatic Speech Emotion Recognition

ICASSP 2018accepted

The scarcity of emotional speech data is a bottleneck of developing automatic speech emotion recognition (ASER) systems. One way to alleviate this issue is to use unsupervised feature learning techniques to learn features from the widely available general speech and use these features to train emoti…

Cited by 0SourceScholar
2018

Visualization and Interpretation of Siamese Style Convolutional Neural Networks for Sound Search by Vocal Imitation

ICASSP 2018accepted

Designing systems that allow users to search sounds through vocal imitation augments the current text-based search engines and advances human-computer interaction. Previously we proposed a Siamese style convolutional network called IMINET for sound search by vocal imitation, which jointly addresses…

Cited by 0SourceScholar
2017

See and listen: Score-informed association of sound tracks to players in chamber music performance videos

ICASSP 2017accepted

Both audio and visual aspects of a musical performance, especially their association, are important for expressing players' ideas and for engaging the audience. In this paper, we present a framework for combining audio and video analyses of multi-instrument chamber music performances to associate pl…

Cited by 0SourceScholar
2017

Visually informed multi-pitch analysis of string ensembles

ICASSP 2017accepted

Multi-pitch analysis of polyphonic music requires estimating concurrent pitches (estimation) and organizing them into temporal streams according to their sound sources (streaming). This is challenging for approaches based on audio alone due to the polyphonic nature of the audio signals. Video of the…

Cited by 0SourceScholar
2016

Emotion classification: How does an automated system compare to Naive human coders?

ICASSP 2016accepted

The fact that emotions play a vital role in social interactions, along with the demand for novel human-computer interaction applications, have led to the development of a number of automatic emotion classification systems. However, it is still debatable whether the performance of such systems can co…

Cited by 0SourceScholar