← Search

Lingyan Huang

2 accepted papers

2024

Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic Attention

ICASSP 2024accepted

End-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’…

Cited by 0SourceScholar
2024

MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis

AAAI 2024technical

The style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, whi…