← Search

Alexandre Costa Ferro Filho

2 accepted papers

2026

TAGARELA - A PORTUGUESE SPEECH DATASET FROM PODCASTS

ICASSP 2026poster

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over 8,972 hours of podcast audio, specifically curated for training…

Cited by 0SourcePDFScholar
2025

BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese Zero-Shot TTS

EMNLP 2025

The detection of audio deepfakes (ADD) has become increasingly important due to the rapid evolution of generative speech models. However, progress in this field remains uneven across languages, particularly for low-resource languages like Portuguese, which lack high-quality datasets. In this paper,