COLING 2024main0 citations

Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS

Gérard Bailly, Romain Legrand, Martin Lenglet, Frédéric Elisei, Maëva Garnier, Olivier Perrotin

Abstract

We developped a web app for ascribing verbal descriptions to expressive audiovisual utterances. These descriptions are limited to lists of adjectives that are either suggested via a navigation in emotional latent spaces built using discriminant analysis of BERT embeddings or entered freely by subjects. We show that such verbal descriptions collected on-line via Prolific on massive data (310 participants, 12620 labelled utterances up-to-now) provide Expressive Multimodal Text-to-Speech Synthesis with precise verbal control over desired emotional content

BibTeX
@inproceedings{bailly-etal-2024-emotags,
    title = "Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal {TTS}",
    author = {Bailly, G{\'e}rard  and
      Legrand, Romain  and
      Lenglet, Martin  and
      Elisei, Fr{\'e}d{\'e}ric  and
      Garnier, Ma{\"e}va  and
      Perrotin, Olivier},
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.505/",
    pages = "5689--5695"
}