← Search

Soheil Khorram

6 accepted papers

2025

Weak-to-Strong Generalization in Speech Recognition

ICASSP 2025accepted

To surpass human-level accuracy, speech recognition models must go beyond relying solely on human labels. To this end, we must build stronger models from weaker supervisors and this is the main goal in weak-to-strong generalization (WSG). WSG methods normally incorporate additional information into…

Cited by 0SourceScholar
2024

Monte Carlo Self-Training for Speech Recognition

ICASSP 2024accepted

Self-training in the teacher-student framework generally suffers from the confirmation bias problem, where errors from the teacher are propagated to the student and hence get amplified with multiple iterations. In this paper, we present Monte Carlo Self-training where pseudo labels are generated by…

Cited by 0SourceScholar
2023

Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition

ICASSP 2023accepted

Semi-supervised training can be performed by jointly optimizing supervised and unsupervised losses. In many settings, supervised and unsupervised losses are inconsistent, and this inconsistency creates instability in training. As a solution, we propose cross-training: instead of training one network…

Cited by 3SourceScholar
2022

Contrastive Siamese Network for Semi-Supervised Speech Recognition

ICASSP 2022accepted

This paper introduces contrastive siamese (c-siam) network, an architecture for leveraging unlabeled acoustic data in speech recognition. c-siam is the first network that extracts high-level linguistic information from speech by matching outputs of two identical transformer encoders. It contains aug…

Cited by 0SourceScholar
2019

Exploiting Acoustic and Lexical Properties of Phonemes to Recognize Valence from Speech

ICASSP 2019accepted

Emotions modulate speech acoustics as well as language. The latter influences the sequences of phonemes that are produced, which in turn further modulate the acoustics. Therefore, phonemes impact emotion recognition in two ways: (1) they introduce an additional source of variability in speech signal…

Cited by 0SourceScholar
2019

Trainable Time Warping: Aligning Time-series in the Continuous-time Domain

ICASSP 2019accepted

DTW calculates the similarity or alignment between two signals, subject to temporal warping. However, its computational complexity grows exponentially with the number of time-series. Although there have been algorithms developed that are linear in the number of time-series, they are generally quadra…

Cited by 0SourceScholar