← Search

Yitao Duan

4 accepted papers

2025

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

ICASSP 2025accepted

Combining face-swapping with lip synchronization offers a cost-effective solution for generating customized talking faces. However, directly cascading existing models can introduce significant interference and reduce video clarity due to limited interaction space in the low-level RGB domain. To solv…

Cited by 0SourceScholar
2024

A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models

COLING 2024main

Machine Translation (MT) has greatly advanced over the years due to the developments in deep neural networks. However, the emergence of Large Language Models (LLMs) like GPT-4 and ChatGPT is introducing a new phase in the MT domain. In this context, we believe that the future of MT is intricately ti…

Cited by 16SourcePDFScholar
2022

Semantically Consistent Data Augmentation for Neural Machine Translation via Conditional Masked Language Model

COLING 2022main

This paper introduces a new data augmentation method for neural machine translation that can enforce stronger semantic consistency both within and across languages. Our method is based on Conditional Masked Language Model (CMLM) which is bi-directional and can be conditional on both left and right c…

2021

An End-to-End Speech Accent Recognition Method Based on Hybrid CTC/Attention Transformer ASR

ICASSP 2021accepted

This paper proposes a novel accent recognition system in the framework of a transformer-based end-to-end speech recognition system. To incorporate the pronunciation and linguistic knowledge into the network, we first pre-train an ASR model in a hybrid CTC/attention manner. Then, focusing on accent r…

Cited by 0SourceScholar