← Search

Lin-shan Lee

15 accepted papers

2024

REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR

NeurIPS 2024poster

Unsupervised automatic speech recognition (ASR) aims to learn the mapping between the speech signal and its corresponding textual transcription without the supervision of paired speech-text data. A word/phoneme in the speech signal is represented by a segment of speech signal with variable length an…

2024

SpeechDPR: End-To-End Spoken Passage Retrieval For Open-Domain Spoken Question Answering

ICASSP 2024accepted

Spoken Question Answering (SQA) is essential for machines to reply to user’s question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domai…

Cited by 0SourceScholar
2021

Fragmentvc: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with Attention

ICASSP 2021accepted

Any-to-any voice conversion aims to convert the voice from and to any speakers even unseen during training, which is much more challenging compared to one-to-one or many-to-many tasks, but much more attractive in real-world scenarios. In this paper we proposed FragmentVC, in which the latent phoneti…

Cited by 0SourceScholar
2020

Interrupted and Cascaded Permutation Invariant Training for Speech Separation

ICASSP 2020accepted

Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the minimum cost label assignments dynamically, very few studies considered the separation problem to be optimizing both the mod…

Cited by 0SourceScholar
2020

Sequence-to-Sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding

ICASSP 2020accepted

In this paper, we investigate the benefit that off-the-shelf word embedding can bring to the sequence-to-sequence (seq-to-seq) automatic speech recognition (ASR). We first introduced the word embedding regularization by maximizing the cosine similarity between a transformed decoder feature and the t…

Cited by 0SourceScholar
2020

Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning

ICASSP 2020accepted

In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is achieved by proper temporal segmentation to make the representat…

Cited by 0SourceScholar
2019

Adversarial Training of End-to-end Speech Recognition Using a Criticizing Language Model

ICASSP 2019accepted

In this paper we proposed a novel Adversarial Training (AT) approach for end-to-end speech recognition using a Criticizing Language Model (CLM). In this way the CLM and the automatic speech recognition (ASR) model can challenge and learn from each other iteratively to improve the performance. Since…

Cited by 0SourceScholar
2019

Towards End-to-end Speech-to-text Translation with Two-pass Decoding

ICASSP 2019accepted

Speech-to-text translation (ST) refers to transforming the audio in source language to the text in target language. Mainstream solutions for such tasks are to cascade automatic speech recognition with machine translation, for which the transcriptions of the source language are needed in training. En…

Cited by 0SourceScholar
2018

Domain Independent Key Term Extraction from Spoken Content Based on Context and Term Location Information in the Utterances

ICASSP 2018accepted

This paper proposes a domain independent approach for extracting key terms from spoken content based on context and term location information, or the sentence structures. Once it is trained with data of enough different domains, it is able to extract key terms in other unseen domains. This is obviou…

Cited by 0SourceScholar
2018

Scalable Sentiment for Sequence-to-Sequence Chatbot Response with Performance Analysis

ICASSP 2018accepted

Conventional seq2seq chatbot models only try to find the sentences with the highest probabilities conditioned on the input sequences, without considering the sentiment of the output sentences. Some research works trying to modify the sentiment of the output sequences were reported. In this paper, we…

Cited by 0SourceScholar
2018

Segmental Audio Word2Vec: Representing Utterances as Sequences of Vectors with Applications in Spoken Term Detection

ICASSP 2018accepted

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be trained in an unsupervised way from an unlabeled corpus, exce…

Cited by 0SourceScholar
2018

Transcribing Lyrics from Commercial Song Audio: the First Step Towards Singing Content Processing

ICASSP 2018accepted

Spoken content processing (such as retrieval and browsing) is maturing, but the singing content is still almost completely left out. Songs are human voice carrying plenty of semantic information just as speech, and may be considered as a special type of speech with highly flexible prosody. The vario…

Cited by 0SourceScholar
2017

Personalized acoustic modeling by weakly supervised multi-task deep learning using acoustic tokens discovered from unlabeled data

ICASSP 2017accepted

It is well known that recognizers personalized to each user are much more effective than user-independent recognizers. With the popularity of smartphones today, although it is not difficult to collect a large set of audio data for each user, it is difficult to transcribe it. However, it is now possi…

Cited by 0SourceScholar
2015

Enhancing automatically discovered multi-level acoustic patterns considering context consistency with applications in spoken term detection

ICASSP 2015accepted

This paper presents a novel approach for enhancing the multiple sets of acoustic patterns automatically discovered from a given corpus. In a previous work it was proposed that different HMM configurations (number of states per model, number of distinct models) for the acoustic patterns form a two-di…

Cited by 0SourceScholar
2015

Enhancing sparse voice annotation for semantic retrieval of personal photos by continuous space word representations

ICASSP 2015accepted

It is very attractive for the user to retrieve photos from a huge collection using high-level personal queries (e.q. uncle Bill's house), but technically very challenging. The previous work proposed a set of approaches to achieve the goal assuming only 30% of the photos are annotated by sparse spoke…

Cited by 0SourceScholar