← Search

Cheol Jun Cho

6 accepted papers

2025

Sylber: Syllabic Embedding Representation of Speech from Raw Audio

ICLR 2025poster

Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequences that are costly to process. To bridge this gap, we propose a new model, Sylbe…

2024

SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in Hubert

ICASSP 2024accepted

Data-driven unit discovery in self-supervised learning (SSL) of speech has embarked on a new era of spoken language processing. Yet, the discovered units often remain in phonetic space and speech units beyond phonemes are largely underexplored. Here, we demonstrate that a syllabic organization emerg…

Cited by 0SourceScholar
2024

Self-Supervised Models of Speech Infer Universal Articulatory Kinematics

ICASSP 2024accepted

Self-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun “probing” models like HuBERT, to correlate their internal representations to different as…

Cited by 0SourceScholar
2023

Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech

ICASSP 2023accepted

Recent self-supervised learning (SSL) models have proven to learn rich representations of speech, which can readily be utilized by diverse downstream tasks. To understand such utilities, various analyses have been done for speech SSL models to reveal which and how information is encoded in the learn…

Cited by 0SourceScholar
2023

Neural Latent Aligner: Cross-trial Alignment for Learning Representations of Complex, Naturalistic Neural Data

ICML 2023poster

Understanding the neural implementation of complex human behaviors is one of the major goals in neuroscience. To this end, it is crucial to find a true representation of the neural data, which is challenging due to the high complexity of behaviors and the low signal-to-ratio (SNR) of the signals. He…

Cited by 8SourcePDFScholar
2023

Speaker-Independent Acoustic-to-Articulatory Speech Inversion

ICASSP 2023accepted

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech…

Cited by 0SourceScholar