← Search

Junkun Chen

6 accepted papers

2022

A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing

ICML 2022spotlight

Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding, but for the inverse direction, speech synthesis, the potential…

2022

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

NAACL 2022system demonstrations

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleS…

2021

Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation

ICML 2021spotlight

Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a unified representation for both speech and text is needed by tasks such as end-…

2021

Improving Simultaneous Translation by Incorporating Pseudo-References with Fewer Reorderings

EMNLP 2021main

Simultaneous translation is vastly different from full-sentence translation, in the sense that it starts translation before the source sentence ends, with only a few words delay. However, due to the lack of large-scale, high-quality simultaneous translation datasets, most such systems are still trai…

Cited by 21SourcePDFScholar
2021

RNNLogic: Learning Logic Rules for Reasoning on Knowledge Graphs

ICLR 2021poster

This paper studies learning logic rules for reasoning on knowledge graphs. Logic rules provide interpretable explanations when used for prediction as well as being able to generalize to other tasks, and hence are critical to learn. Existing methods either suffer from the problem of searching in a la…

2019

VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research

ICCV 2019oral

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSR-VTT dataset, \vatex i…

Cited by 669PDFScholar