← Search

Xinhui Li

12 accepted papers

2025

Conditional Information Bottleneck-Based Multivariate Time Series Forecasting

IJCAI 2025

Multivariate time series (MTS) forecasting endeavors to anticipate the forthcoming sequence of interdependent variables through the utilization of past observations. The prevailing methodologies, relying on deep neural networks, Transformer, or information bottleneck frameworks, persist in confronti

2025

EEG Correlation Analysis-guided Graph Local Enhanced Feature Learning For Emotion Recognition

ICASSP 2025accepted

EEG-based emotion recognition is a key technology in brain-computer interfaces. Many previous studies have applied deep learning methods to mine emotion-related features in EEG to decode emotions. However, they overlooked the importance of electrode correlations and varying brain region activation d…

Cited by 0SourceScholar
2025

Probabilistic Semantics Guided Discovery of Approximate Functional Dependencies

UAI 2025

As the general description of relationships between attributes, approximate functional dependencies (AFDs) almost hold for a given dataset with a few violations. Most of existing methods for AFD discover are insufficient to balance the efficiency and accuracy due to the massive search space and perm

2023

Adaptive Texture Filtering for Single-Domain Generalized Segmentation

AAAI 2023technical

Domain generalization in semantic segmentation aims to alleviate the performance degradation on unseen domains through learning domain-invariant features. Existing methods diversify images in the source domain by adding complex or even abnormal textures to reduce the sensitivity to domain-specific f…

Cited by 7SourcePDFScholar
2023

WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training

ICASSP 2023accepted

This paper aims to introduce a robust singing voice synthesis (SVS) system to produce very natural and realistic singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random area conditional discriminators to help supervise the acoust…

Cited by 0SourceScholar
2022

Zero-Shot Cross-Lingual Transfer Using Multi-Stream Encoder and Efficient Speaker Representation

ICASSP 2022accepted

We propose a novel method for zero-shot cross-lingual TTS task by using multi-stream text encoder and efficient speaker representation. Specifically, a unified multi-stream text encoder that takes both advantages of Transformer and CBHG is proposed to retain multiple hypotheses about input represent…

Cited by 0SourceScholar
2021

A New High Quality Trajectory Tiling Based Hybrid TTS In Real Time

ICASSP 2021accepted

A trajectory tiling based, hybrid TTS is revisited in this study for improving its synthesis performance. A combination of Transformer encoder and RNN based decoder architecture where two-level, at both word and Chinese phonetic alphabet letter levels, linguistic representation is exploited to gener…

Cited by 0SourceScholar
2021

Investigation of Fast and Efficient Methods for Multi-Speaker Modeling and Speaker Adaptation

ICASSP 2021accepted

In this paper, we propose a novel method for fast and efficient few-shot TTS task, which is able to disentangle linguistic and speaker representations. Specifically, an adversarial training strategy is firstly employed to wipe out speaker information from the linguistic representations. Then the spe…

Cited by 0SourceScholar
2020

An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training Data

ICASSP 2020accepted

A frame-unit-selection based voice conversion system proposed earlier by us is revisited here to enhance its performance in both speech naturalness and speaker similarity. Speaker independent, bilingual (Mandarin Chinese and American English) deep neural net (DNN) acoustic model’s output, frame-leve…

Cited by 1SourceScholar
2020

Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer

ICASSP 2020accepted

Although Transformer based neural end-to-end TTS model has demonstrated extreme effectiveness in capturing long-term dependencies and achieved state-of-the-art performance, it still suffers from two problems. 1) limited ability to model sequential and local structures in sequences; 2) heavily rely o…

Cited by 0SourceScholar