← Search

Pirros Tsiakoulis

5 accepted papers

2024

Improved Text Emotion Prediction Using Combined Valence and Arousal Ordinal Classification

NAACL 2024short

Emotion detection in textual data has received growing interest in recent years, as it is pivotal for developing empathetic human-computer interaction systems.This paper introduces a method for categorizing emotions from text, which acknowledges and differentiates between the diversified similaritie…

Cited by 6SourcePDFScholar
2023

Investigating Content-Aware Neural Text-to-Speech MOS Prediction Using Prosodic and Linguistic Features

ICASSP 2023accepted

Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a pretrained selfsupervised learning model that directly uses the…

Cited by 0SourceScholar
2021

Prosodic Clustering for Phoneme-Level Prosody Control in End-to-End Speech Synthesis

ICASSP 2021accepted

This paper presents a method for controlling the prosody at the phoneme level in an autoregressive attention-based text-to-speech system. Instead of learning latent prosodic features with a variational framework as is commonly done, we directly extract phoneme-level F0 and duration features from the…

Cited by 12SourceScholar
2015

Distributed dialogue policies for multi-domain statistical dialogue management

ICASSP 2015accepted

Statistical dialogue systems offer the potential to reduce costs by learning policies automatically on-line, but are not designed to scale to large open-domains. This paper proposes a hierarchical distributed dialogue architecture in which policies are organised in a class hierarchy aligned to an un…

Cited by 0SourceScholar
2015

Improving multiple-crowd-sourced transcriptions using a speech recogniser

ICASSP 2015accepted

This paper introduces a method to produce high-quality transcriptions of speech data from only two crowd-sourced transcriptions. These transcriptions, produced cheaply by people on the Internet, for example through Amazon Mechanical Turk, are often of low quality. Often, multiple crowd-sourced trans…

Cited by 23SourceScholar