Speaker Anonymization for Children
Speaker anonymization aims to modify the speech signal in order to protect the identity of a speaker while preserving the linguistic content. Despite the increasing use of children
8 accepted papers
Speaker anonymization aims to modify the speech signal in order to protect the identity of a speaker while preserving the linguistic content. Despite the increasing use of children
Reading fluency in any language requires accurate word decoding but also natural prosodic phrasing i.e the grouping of words into rhythmically and syntactically coherent units. This holds for, both, reading aloud and silent reading. While adults pause meaningfully at clause or punctuation boundaries
Reliable assessment of oral reading fluency (ORF) is of great importance in foundational literacy missions globally. For the design of level appropriate testing passages, text difficulty has traditionally been based on coarse-grained measures of readability like the Flesch–Kincaid score. We present…
The detection of perceived prominence in speech has attracted approaches ranging from the design of knowledge-based linguistic and acoustic features to the automatic feature learning from suprasegmental attributes such as pitch and intensity contours. We present here, in contrast, a system that oper…
In this paper our goal is to convert a set of spoken lines into sung ones. Unlike previous signal processing based methods, we take a learning based approach to the problem. This allows us to automatically model various aspects of this transformation, thus overcoming dependence on specific inputs su…
With the advent of data-driven statistical modeling and abundant computing power, researchers are turning increasingly to deep learning for audio synthesis. These methods try to model audio signals directly in the time or frequency domain. In the interest of more flexible control over the generated…
In this paper, various regularizations on the room impulse response (RIR) are proposed to obtain better single-channel speech dereverberation in the non-negative matrix factorization (NMF) framework. The regularizations on the RIR are motivated by the spectral domain representation of the RIR. To ob…
Structural segmentation of music involves identifying boundaries between homogenous regions where the homogeneity involves one or more musical dimensions, and therefore depends on the musical genre. In this work, we address the segmentation of Hindustani instrumental concert recordings at the highes…