← Search

Milos Cernak

14 accepted papers

2026

STREAMMARK: A DEEP LEARNING-BASED SEMI-FRAGILE AUDIO WATERMARKING FOR PROACTIVE DEEPFAKE DETECTION

ICASSP 2026poster

The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive detection methods, we propose StreamMark, a novel deep learning-based, semi-fragile audio watermarking system. StreamMa…

Cited by 0SourcePDFScholar
2025

Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task

ICASSP 2025accepted

Human perception has the unique ability to focus on specific events in a mixture of signals – a challenging task for existing non-intrusive assessment methods. In this work, we introduce semi-intrusive assessment that emulates human attention by framing audio assessment as a text-prediction task wit…

Cited by 0SourceScholar
2024

Multi-Channel Mosra: Mean Opinion Score and Room Acoustics Estimation Using Simulated Data and A Teacher Model

ICASSP 2024accepted

Previous methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. However, quality-based device selection for rooms with multiple recording device…

Cited by 0SourceScholar
2024

On Real-Time Multi-Stage Speech Enhancement Systems

ICASSP 2024accepted

Recently, multi-stage systems have stood out among deep learning-based speech enhancement methods. However, these systems are always high in complexity, requiring millions of parameters and powerful computational resources, which limits their application for real-time processing in low-power devices…

Cited by 0SourceScholar
2023

Efficient Speech Quality Assessment Using Self-Supervised Framewise Embeddings

ICASSP 2023accepted

Automatic speech quality assessment is essential for audio researchers, developers, speech and language pathologists, and system quality engineers. The current state-of-the-art systems are based on framewise speech features (hand-engineered or learnable) combined with time dependency modeling. This…

Cited by 0SourceScholar
2023

Personalized Task Load Prediction in Speech Communication

ICASSP 2023accepted

Estimating the quality of remote speech communication is a complex task influenced by the speaker, transmission channel, and listener. For example, the degradation of transmission quality can increase listeners’ cognitive load, which can influence the overall perceived quality of the conversation. T…

Cited by 0SourceScholar
2022

SERAB: A Multi-Lingual Benchmark for Speech Emotion Recognition

ICASSP 2022accepted

Recent developments in speech emotion recognition (SER) often leverage deep neural networks (DNNs). Comparing and benchmarking different DNN models can often be tedious due to the use of different datasets and evaluation protocols. To facilitate the process, here, we present the Speech Emotion Recog…

Cited by 0SourceScholar
2020

A Bin Encoding Training of a Spiking Neural Network Based Voice Activity Detection

ICASSP 2020accepted

Advances of deep learning for Artificial Neural Networks (ANNs) have led to significant improvements in the performance of digital signal processing systems implemented on digital chips. Although recent progress in low-power chips is remarkable, neuromorphic chips that run Spiking Neural Networks (S…

Cited by 0SourceScholar
2020

Spiking Neural Networks Trained With Backpropagation for Low Power Neuromorphic Implementation of Voice Activity Detection

ICASSP 2020accepted

Recent advances in Voice Activity Detection (VAD) are driven by artificial and Recurrent Neural Networks (RNNs), however, using a VAD system in battery-operated devices requires further power efficiency. This can be achieved by neuromorphic hardware, which enables Spiking Neural Networks (SNNs) to p…

Cited by 0SourceScholar
2017

Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease

ICASSP 2017accepted

Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized…

Cited by 35SourceScholar
2017

On the impact of non-modal phonation on phonological features

ICASSP 2017accepted

Different modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-mod…

Cited by 0SourceScholar