← Search

Shabnam Ghaffarzadegan

8 accepted papers

2024

CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision

ICASSP 2024accepted

Speech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agn…

Cited by 6SourceScholar
2024

Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?

ICASSP 2024accepted

In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not b…

Cited by 13SourceScholar
2021

Unsupervised Discriminative Learning of Sounds for Audio Event Classification

ICASSP 2021accepted

Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale visual datasets is time consuming. On several audio event c…

Cited by 5SourceScholar
2016

Exploring deep learning architectures for automatically grading non-native spontaneous speech

ICASSP 2016accepted

We investigate two deep learning architectures reported to have superior performance in ASR over the conventional GMM system, with respect to automatic speech scoring. We use an approximately 800-hour large-vocabulary non-native spontaneous English corpus to build three ASR systems. One system is in…

Cited by 0SourceScholar
2015

Generative modeling of pseudo-target domain adaptation samples for whispered speech recognition

ICASSP 2015accepted

The lack of available large corpora of transcribed whispered speech is one of the major roadblocks for development of successful whisper recognition engines. Our recent study has introduced a Vector Taylor Series (VTS) approach to pseudo-whisper sample generation which requires availability of only…

Cited by 0SourceScholar
2015

Leveraging automatic speech recognition in cochlear implants for improved speech intelligibility under reverberation

ICASSP 2015accepted

Despite recent advancements in digital signal processing technology for cochlear implant (CI) devices, there still remains a significant gap between speech identification performance of CI users in reverberation compared to that in anechoic quiet conditions. Alternatively, automatic speech recogniti…

Cited by 0SourceScholar