← Search

Cheng Gong

12 accepted papers

2025

Cross-Scenario End-to-End Motion Planning in Off-Road Environment: A Lifelong Learning Perspective

RA-L 2025

Motion planning in off-road scenarios is particularly challenging due to diverse terrain features, surface characteristics, and environmental factors. Consequently, rule-based or fixed-parameter motion planning methods often fail to maintain optimal performance, especially in cross-scenario applicat

Cited by 2SourceScholar
2025

Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge

ICASSP 2025accepted

In this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we pr…

Cited by 0SourceScholar
2025

MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework

CVPR 2025poster

Crafting adversarial examples is crucial for evaluating and enhancing the robustness of Deep Neural Networks (DNNs), presenting a challenge equivalent to maximizing a non-differentiable 0-1 loss function. However, existing single objective methods, namely adversarial attacks focus on a surrogate…

2025

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

NeurIPS 2025spotlight

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modelin…

Cited by 0SourceScholar
2024

Deep Feature Surgery: Towards Accurate and Efficient Multi-Exit Networks

ECCV 2024poster

"Multi-exit network is a promising architecture for efficient model inference by sharing backbone networks and weights among multiple exits. However, the gradient conflict of the shared weights results in sub-optimal accuracy. This paper introduces Deep Feature Surgery (), which consists of feature…

2024

Learning Pareto Set for Multi-Objective Continuous Robot Control

IJCAI 2024poster

For a control problem with multiple conflicting objectives, there exists a set of Pareto-optimal policies called the Pareto set instead of a single optimal policy. When a multi-objective control problem is continuous and complex, traditional multi-objective reinforcement learning (MORL) algorithms s…

2023

VF-Taco2: Towards Fast and Lightweight Synthesis for Autoregressive Models with Variation Autoencoder and Feature Distillation

ICASSP 2023accepted

With the development of deep learning, end-to-end neural text-to-speech (TTS) systems have achieved significant improvements in high-quality speech synthesis. However, most of these systems are attention-based autoregressive models, resulting in slow synthesis speed and large model parameter sizes.…

Cited by 0SourceScholar
2022

Joint and Adversarial Training with ASR for Expressive Speech Synthesis

ICASSP 2022accepted

Style modeling is an important issue and has been proposed in expressive speech synthesis. In existing unsupervised methods, the style encoder extracts the latent representation from the reference audio as style information. However, the style information extracted from the style encoder will entang…

Cited by 0SourceScholar
2022

Using Multiple Reference Audios and Style Embedding Constraints for Speech Synthesis

ICASSP 2022accepted

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding must be manually selected during inference. Due to the fact t…

Cited by 0SourceScholar
2021

Didispeech: A Large Scale Mandarin Speech Corpus

ICASSP 2021accepted

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is recorded in quiet environment and is suitable for various speech p…

Cited by 0SourceScholar
2021

Improving Naturalness and Controllability of Sequence-to-Sequence Speech Synthesis by Learning Local Prosody Representations

ICASSP 2021accepted

State-of-the-art neural text-to-speech (TTS) networks are trained with a large amount of speech data, which significantly improves the quality of synthetic speech compared with traditional approaches. However, the prosody and controllability of the generated speech is still insufficient, especially…

Cited by 0SourceScholar
2021

Orientation-Aware Planning for Parallel Task Execution of Omni-Directional Mobile Robot

IROS 2021poster

Omni-directional mobile robot (OMR) systems have been very popular in academia and industry for their superb maneuverability and flexibility. Yet their potential has not been fully exploited, where the extra degree of freedom in OMR can potentially enable the robot to carry out extra tasks. For inst…

Cited by 2SourceScholar