← Search

Max Morrison

4 accepted papers

2024

Crowdsourced and Automatic Speech Prominence Estimation

ICASSP 2024accepted

The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the prominence of each word in an utterance. These prominence labels are…

Cited by 0SourceScholar
2022

Chunked Autoregressive GAN for Conditional Waveform Synthesis

ICLR 2022poster

Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI. These systems employ deep generative models that model the waveform via either sequential (autoregressive) or parallel (non-autoregressive) sampling. Generative a…

2022

VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio Models

NeurIPS 2022accept

As governments and corporations adopt deep learning systems to collect and analyze user-generated audio data, concerns about security and privacy naturally emerge in areas such as automatic speaker recognition. While audio adversarial examples offer one route to mislead or evade these invasive syste…

2021

Context-Aware Prosody Correction for Text-Based Speech Editing

ICASSP 2021accepted

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings often sound unnatural because of prosody mismatches around e…

Cited by 0SourceScholar