← Search

Okko Räsänen

6 accepted papers

2022

Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases

ICASSP 2022accepted

We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated audio clips and their captions by scoring the similarities b…

Cited by 0SourceScholar
2021

Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections

ICASSP 2021accepted

In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Zero-shot learning in audio classification refers to classification problems that aim at recognizing audio instances of so…

Cited by 0SourceScholar
2019

Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style Conversion

ICASSP 2019accepted

Speaking style conversion (SSC) is the technology of converting natural speech signals from one style to another. In this study, we propose the use of cycle-consistent adversarial networks (CycleGANs) for converting styles with varying vocal effort, and focus on conversion between normal and Lombard…

Cited by 0SourceScholar
2019

Data Augmentation Strategies for Neural Network F0 Estimation

ICASSP 2019accepted

This study explores various speech data augmentation methods for the task of noise-robust fundamental frequency (F0) estimation with neural networks. The explored augmentation strategies are split into additive noise and channel-based augmentation and into vocoder-based augmentation methods. In voco…

Cited by 0SourceScholar