2026
TAGARELA - A PORTUGUESE SPEECH DATASET FROM PODCASTS
Frederico Santos de Oliveira, Lucas Rafael Stefanel Gris, Augusto Seben da Rosa, Alexandre Costa Ferro Filho, Edresson Casanova, Christopher Dane Shulby +2
ICASSP 2026poster
Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over 8,972 hours of podcast audio, specifically curated for training…