← Search

Xinjian Li

12 accepted papers

2026

Spontaneous Yet Predictable: Shapelet-Driven, Channel-Aware Intention Decoding from Multi-Region ECoG

AAAI 2026technical

Proactive intention decoding remains a critical yet underexplored challenge in brain–machine interfaces (BMIs), especially under naturalistic, self-initiated behavior. Existing systems rely on reactive decoding of motor cortex signals, resulting in substantial latency. To address this, we leverage t

Cited by 0SourcePDFScholar
2024

Towards Robust Speech Representation Learning for Thousands of Languages

EMNLP 2024main

Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world’s 7000+ languages. We propose XEUS, a Cross-lingual Encoder for Universal Speech, trained on over 1 million hours of d…

2023

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

IJCAI 2023poster

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the…

2023

Textless Direct Speech-to-Speech Translation with Discrete Speech Representation

ICASSP 2023accepted

Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade systems, which are often composed of recognition, translation and synthesis sub-systems. However, most of end-to-end system…

Cited by 0SourceScholar
2022

On Adversarial Robustness Of Large-Scale Audio Visual Learning

ICASSP 2022accepted

As audio-visual systems are being deployed for safety-critical tasks such as surveillance and malicious content filtering, their robustness remains an under-studied area. Existing published work on robustness either does not scale to large-scale dataset, or does not deal with multiple modalities. Th…

Cited by 0SourceScholar
2022

Zero-shot Learning for Grapheme to Phoneme Conversion with Language Ensemble

ACL 2022findings

Grapheme-to-Phoneme (G2P) has many applications in NLP and speech fields. Most existing work focuses heavily on languages with abundant training datasets, which limits the scope of target languages to less than 100 languages. This work attempts to apply zero-shot learning to approximate G2P models f…

2021

Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource Languages

ICASSP 2021accepted

With recent advancements in language technologies, humans are now speaking to devices. Increasing the reach of spoken language technologies requires building systems in local languages. A major bottleneck here are the underlying data-intensive parts that make up such systems, including automatic spe…

Cited by 0SourceScholar
2021

Multilingual Phonetic Dataset for Low Resource Speech Recognition

ICASSP 2021accepted

Phone Recognition is one of the most important tasks in the field of multilingual speech recognition, especially for low-resource languages whose orthographies are not available. However, most speech recognition datasets so far only focus on high-resource languages, there are very few datasets avail…

Cited by 0SourceScholar
2021

Phone Distribution Estimation for Low Resource Languages

ICASSP 2021accepted

Phones are critical components in various computational linguistic fields, for example, phone distributions could be helpful in speech recognition and speech synthesis. Traditional approaches to estimate phone distributions typically involve G2P systems which are either manually designed by linguist…

Cited by 0SourceScholar
2020

Universal Phone Recognition with a Multilingual Allophone System

ICASSP 2020accepted

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and thei…

Cited by 0SourceScholar
2019

Adversarial Music: Real world Audio Adversary against Wake-word Detection System

NeurIPS 2019spotlight

Voice Assistants (VAs) such as Amazon Alexa or Google Assistant rely on wake-word detection to respond to people's commands, which could potentially be vulnerable to audio adversarial examples. In this work, we target our attack on the wake-word detection system. Our goal is to jam the model with so…

Cited by 73SourcePDFScholar
2019

Phoneme Level Language Models for Sequence Based Low Resource ASR

ICASSP 2019accepted

Building multilingual and crosslingual models help bring different languages together in a language universal space. It allows models to share parameters and transfer knowledge across languages, enabling faster and better adaptation to a new language. These approaches are particularly useful for low…

Cited by 0SourceScholar