← Search

Pengcheng Zhu

12 accepted papers

2026

InstGPMap: Real-Time Instance-Level Global Prior Mapping via Historical Predictions Fusion

RA-L 2026

Recent online methods for HD map construction directly infer local maps from sensor observations, yet suffer from limited perception range, particularly under challenging scenarios such as occlusions by large vehicles or poor visibility in rainy conditions. Inspired by human perception, which increm

Cited by 0SourceScholar
2026

MEANVC: LIGHTWEIGHT AND STREAMING ZERO-SHOT VOICE CONVERSION VIA MEAN FLOWS

ICASSP 2026poster

Zero-shot voice conversion (VC) aims to transfer timbre from a source speaker to any unseen target speaker while preserving linguistic content. Growing application scenarios demand models with streaming inference capabilities. This has created a pressing need for models that are simultaneously fast,…

Cited by 0SourcePDFScholar
2025

BEV-PolyNet: BEV-Based Polygonal End to End Parking Slot Detection Framework

RA-L 2025

Accurate parking slot detection is crucial for autonomous parking and intelligent driving, directly impacting safety and efficiency. However, most existing methods rely on AVM images, making them susceptible to image distortion and vehicle occlusion. Additionally, many approaches independently regre

Cited by 0SourceScholar
2025

MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion

ICASSP 2025accepted

In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same…

Cited by 0SourceScholar
2025

Robust Supervised Graph Embedding Method For EEG-Based Brain Network Emotion Recognition

ICASSP 2025accepted

Emotion recognition based on brain networks has attracted increasing research attention due to its ability to reveal the information interactions between brain regions under different emotional states. However, there are still two challenges in practical applications: 1) The high dimensionality of b…

Cited by 0SourceScholar
2024

Dualvc 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion

ICASSP 2024accepted

Voice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective through streaming model architecture design and intra-model knowledge distillation al…

Cited by 0SourceScholar
2024

MGS-SLAM: Monocular Sparse Tracking and Gaussian Mapping With Depth Smooth Regularization

RA-L 2024

This letter introduces a novel framework for dense Visual Simultaneous Localization and Mapping (VSLAM) based on Gaussian Splatting. Recently, SLAM based on Gaussian Splatting has shown promising results. However, in monocular scenarios, the Gaussian maps reconstructed lack geometric accuracy and ex

Cited by 22SourceScholar
2023

Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features

ICASSP 2023accepted

Voice conversion for highly expressive speech is challenging. Current approaches struggle with the balance between speaker similarity, intelligibility, and expressiveness. To address this problem, we propose Expressive-VC, a novel end-to-end voice conversion framework that leverages advantages from…

Cited by 0SourceScholar
2023

Robust Learning for Multi-party Addressee Recognition with Discrete Addressee Codebook

ACL 2023short

Addressee recognition aims to identify addressees in multi-party conversations. While state-of-the-art addressee recognition models have achieved promising performance, they still suffer from the issue of robustness when applied in real-world scenes. When exposed to a noisy environment, these models…

Cited by 1SourcePDFScholar
2022

One-Shot Voice Conversion For Style Transfer Based On Speaker Adaptation

ICASSP 2022accepted

One-shot style transfer is a challenging task, since training on one utterance makes model extremely easy to over-fit to training data and causes low speaker similarity and lack of expressiveness. In this paper, we build on the recognition-synthesis framework and propose a one-shot voice conversion…

Cited by 0SourceScholar
2022

VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis

ICASSP 2022accepted

In this paper, we propose VISinger, a complete end-to-end high-quality singing voice synthesis (SVS) system that directly generates singing audio from lyrics and musical score. Our approach is inspired by VITS [1], an end-to-end speech generation model which adopts VAE-based posterior encoder augmen…

Cited by 0SourceScholar