← Search

Rao Ma

10 accepted papers

2025

Cross-Lingual Transfer Learning for Speech Translation

NAACL 2025short

There has been increasing interest in building multilingual foundation models for NLP and speech research. This paper examines how to expand the speech translation capability of these models with restricted data. Whisper, a speech foundation model with strong performance on speech recognition and En…

Cited by 1SourcePDFScholar
2025

LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

EMNLP 2025

Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language processing tasks, including Automatic Speech Recognition (ASR). To effectively combine both models for better performance, cont

Cited by 0SourcePDFScholar
2025

Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs

EMNLP 2025

The combination of pre-trained speech encoders with large language models has enabled the development of speech LLMs that can handle a wide range of spoken language processing tasks. While these models are powerful and flexible, this very flexibility may make them more vulnerable to adversarial atta

Cited by 0SourcePDFScholar
2024

Investigating the Emergent Audio Classification Ability of ASR Foundation Models

NAACL 2024long

Text and vision foundation models can perform many tasks in a zero-shot setting, a desirable property that enables these systems to be applied in general and low-resource settings. There has been far less work, however, on the zero-shot abilities of ASR foundation models, with these systems typicall…

2024

Towards End-to-End Spoken Grammatical Error Correction

ICASSP 2024accepted

Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, wi…

Cited by 0SourceScholar
2023

Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain Adaptation

ICASSP 2023accepted

ASR model deployment environment is ever-changing, and the incoming speech can be switched across different domains during a session. This brings a challenge for effective domain adaptation when only target domain text data is available, and our objective is to obtain obviously improved performance…

Cited by 0SourceScholar
2021

AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge

ICASSP 2021accepted

This paper describes the AISpeech-SJTU ASR system for the Interspeech-2020 Accented English Speech Recognition Challenge (AESRC). This task is challenging due to the diversity of pronunciation accuracy, intonation speed and pronunciation of some syllables. All participants were restricted to develop…

Cited by 0SourceScholar
2021

AISpeech-SJTU Accent Identification System for the Accented English Speech Recognition Challenge

ICASSP 2021accepted

This paper describes the AISpeech-SJTU system for the accent identification track of the Interspeech-2020 Accented English Speech Recognition Challenge. In this challenge track, only 160-hour accented English data collected from 8 countries and the auxiliary Librispeech dataset are provided for trai…

Cited by 0SourceScholar
2020

Addressing the Polysemy Problem in Language Modeling with Attentional Multi-Sense Embeddings

ICASSP 2020accepted

Neural network language models have gained considerable popularity due to their promising performance. Distributed word embeddings are utilized to represent semantic information. However, each word is associated with a single vector in the embedding layer, disabling the model from capturing the mean…

Cited by 0SourceScholar