← Search

Shuju Shi

3 accepted papers

2025

What You Read Isn’t What You Hear: Linguistic Sensitivity in Deepfake Speech Detection

EMNLP 2025

Recent advances in text-to-speech technology have enabled highly realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio anti-spoofing systems are critical for detecting such threats, prior research has predominantly focused on acoustic-level per

2023

An ASR-Free Fluency Scoring Approach with Self-Supervised Learning

ICASSP 2023accepted

A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for the subsequent calculation of fluency-related features or directly modeling speech fluency with an end-to-end approach. This paper describes a novel ASR-free ap…

Cited by 0SourceScholar
2023

Leveraging Phone-Level Linguistic-Acoustic Similarity For Utterance-Level Pronunciation Scoring

ICASSP 2023accepted

Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation qu…

Cited by 0SourceScholar