← Search

György Fazekas

14 accepted papers

2025

Leave-One-EquiVariant: Alleviating Invariance-Related Information Loss in Contrastive Music Representations

ICASSP 2025accepted

Contrastive learning has proven effective in self-supervised musical representation learning, particularly for Music Information Retrieval (MIR) tasks. However, reliance on augmentation chains for contrastive view generation and the resulting learnt invariances pose challenges when different downstr…

Cited by 0SourceScholar
2025

Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding

ICASSP 2025accepted

Efficiently compressing high-dimensional audio signals into a compact and informative latent space is crucial for various tasks, including generative modeling and music information retrieval (MIR). Existing audio autoencoders, however, often struggle to achieve high compression ratios while preservi…

Cited by 0SourceScholar
2025

Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores

ICASSP 2025accepted

This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performanc…

Cited by 0SourceScholar
2023

Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution

ICASSP 2023accepted

Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resol…

Cited by 0SourceScholar
2023

HIPI: A Hierarchical Performer Identification Model Based on Symbolic Representation of Music

ICASSP 2023accepted

Automatic Performer Identification from the symbolic representation of music has been a challenging topic in Music Information Retrieval (MIR).In this study, we apply a Recurrent Neural Network (RNN) model to classify the most likely music performers from their interpretative styles. We study differ…

Cited by 0SourceScholar
2023

Rigid-Body Sound Synthesis with Differentiable Modal Resonators

ICASSP 2023accepted

Physical models of rigid bodies are used for sound synthesis in applications from virtual environments to music production. Traditional methods, such as modal synthesis, often rely on computationally expensive numerical solvers, while recent deep learning approaches are limited by post-processing of…

Cited by 0SourceScholar
2022

Learning Music Audio Representations Via Weak Language Supervision

ICASSP 2022accepted

Audio representations for music information retrieval are typically learned via supervised learning in a task-specific fashion. Although effective at producing state-of-the-art results, this scheme lacks flexibility with respect to the range of applications a model can have and requires extensively…

Cited by 0SourceScholar
2022

Violinist Identification Using Note-Level Timbre Feature Distributions

ICASSP 2022accepted

Modelling musical performers’ individual playing styles based on audio features is important for music education, music expression analysis and music generation. In violin performance, the perception of playing styles are mainly affected by the characteristic musical timbre, which is mostly determin…

Cited by 0SourceScholar
2018

Feature Design Using Audio Decomposition for Intelligent Control of the Dynamic Range Compressor

ICASSP 2018accepted

This papeper proposes a method of controlling the dynamic range compressor using sound examples. Our earlier work showed the effectiveness of random forest regression to map acoustic features to effect control parameters [1]. We extend this work to address the challenging task of extracting relevant…

Cited by 0SourceScholar
2017

Convolutional recurrent neural networks for music classification

ICASSP 2017accepted

We introduce a convolutional recurrent neural network (CRNN) for music tagging. CRNNs take advantage of convolutional neural networks (CNNs) for local feature extraction and recurrent neural networks for temporal summarisation of the extracted features. We compare CRNN with three CNN structures that…

Cited by 0SourceScholar
2015

On the use of the tempogram to describe audio content and its application to Music structural segmentation

ICASSP 2015accepted

This paper presents a new set of audio features to describe music content based on tempo cues. Tempogram, a mid-level representation of tempo information, is constructed to characterize tempo variation and local pulse in the audio signal. We introduce a collection of novel tempogram-based features i…

Cited by 0SourceScholar