← Search

Joshua D. Reiss

8 accepted papers

2024

Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley Synthesis

ICASSP 2024accepted

Sound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio with video. In some cases, environmental recordings from video…

Cited by 0SourceScholar
2023

Cross-Modal Fusion Techniques for Utterance-Level Emotion Recognition from Text and Speech

ICASSP 2023accepted

Multimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are particularly important for a human participant in understanding emotions. Although…

Cited by 0SourceScholar
2023

Modelling Black-Box Audio Effects with Time-Varying Feature Modulation

ICASSP 2023accepted

Deep learning approaches for black-box modelling of audio effects have shown promise, however, the majority of existing work focuses on nonlinear effects with behaviour on relatively short time-scales, such as guitar amplifiers and distortion. While recurrent and convolutional architectures can theo…

Cited by 24SourceScholar
2022

Direct Design of Biquad Filter Cascades with Deep Learning by Sampling Random Polynomials

ICASSP 2022accepted

Designing infinite impulse response filters to match an arbitrary magnitude response requires specialized techniques. Methods like modified Yule-Walker are relatively efficient, but may not be sufficiently accurate in matching high order responses. On the other hand, iterative optimization technique…

Cited by 0SourceScholar
2020

Modeling Plate and Spring Reverberation Using A DSP-Informed Deep Neural Network

ICASSP 2020accepted

Plate and spring reverberators are electromechanical systems first used and researched as means to substitute real room reverberation. Currently, they are often used in music production for aesthetic reasons due to their particular sonic characteristics. The modeling of these audio processors and th…

Cited by 0SourceScholar
2019

End-to-End Probabilistic Inference for Nonstationary Audio Analysis

ICML 2019oral

A typical audio signal processing pipeline includes multiple disjoint analysis stages, including calculation of a time-frequency representation followed by spectrogram-based feature analysis. We show how time-frequency analysis and nonnegative matrix factorisation can be jointly formulated as a spec…

Cited by 10SourcePDFScholar
2019

Unifying Probabilistic Models for Time-frequency Analysis

ICASSP 2019accepted

In audio signal processing, probabilistic time-frequency models have many benefits over their non-probabilistic counterparts. They adapt to the incoming signal, quantify uncertainty, and measure correlation between the signal’s amplitude and phase information, making time domain resynthesis straight…

Cited by 0SourceScholar