← Search

Jianguo Wei

29 accepted papers

2026

EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment

AAAI 2026technical

Dysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparit

Cited by 0SourcePDFScholar
2025

Attribute Association Driven Multi-Task Learning for Session-based Recommendation

IJCAI 2025

Session-based Recommendation (SBR) aims to predict users’ next interaction based on their current session without relying on long-term profiles. Despite its effectiveness in privacy-preserving and real-time scenarios, SBR remains challenging due to limited behavioral signals. Prior methods often ove

Cited by 0SourcePDFScholar
2025

Continual Unsupervised Domain Adaptation for Audio Deepfake Detection

ICASSP 2025accepted

Audio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models…

Cited by 0SourceScholar
2025

Dynamic Neighborhood Modeling via Node-Subgraph Contrastive Learning for Graph-Based Fraud Detection

AAAI 2025technical

Fraud detection that aims to discern frauds from the majority of benigns has become an increasingly prominent research field. Recently, Graph Neural Networks (GNNs) have been widely applied in graph-based fraud detection due to their outstanding data analysis and mining capabilities. However, owing…

Cited by 0SourcePDFScholar
2024

EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention Mechanism

ICASSP 2024accepted

Auditory attention detection (AAD) based on electroencephalogram (EEG) helps recognize the target speaker in a cocktail party scenario, advancing auditory brain-computer interface development. Previous EEG studies on AAD were largely based on data collected in laboratory settings. In this study, we…

Cited by 0SourceScholar
2024

Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory Data

ICASSP 2024accepted

Ultrasonic imaging is one of the most popular methods for tracking tongue motion. Imaging plane shift and contact variation are crucial factors affecting the consistency of the obtained ultrasonic images. To solve this issue, researchers proposed many different helmets. In this study, we propose an…

Cited by 0SourceScholar
2024

Generalized Taxonomy-Guided Graph Neural Networks

IJCAI 2024poster

Graph neural networks have been demonstrated to be effective analytic apparatus for mining network data. Most real-world networks are inherently hierarchical, offering unique opportunities to acquire latent, intrinsic network organizational properties by utilizing network taxonomies. The existing ap…

Cited by 0SourcePDFScholar
2024

SSR-GPCsT: Deep Learning Models Based on Functional Connectivity Maps in Autism Research

ICASSP 2024accepted

Autism is a neurodevelopmental disorder characterized by difficulties in social interaction, communication, and sensory sensitivity. Functional magnetic resonance imaging (fMRI) is a commonly used brain imaging technique to obtain functional connectivity information in individuals with autism. Howev…

Cited by 0SourceScholar
2024

Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion Recognition

ICASSP 2024accepted

In the tasks of domain adaptation (DA) for speech emotion recognition (SER), self-supervised learning (SSL) algorithms could effectively explore domain and structural information from target domain samples, thereby mitigating domain discrepancies. However, in a general setting, when the target domai…

Cited by 0SourceScholar
2024

Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image

NeurIPS 2024poster

In the visual spatial understanding (VSU) field, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods for standalone SI2T or ST2I perform imperfectly in spatial understanding, due to the difficulty of 3D-wise spatial featu…

Cited by 0SourcePDFScholar
2023

Generating Visual Spatial Description via Holistic 3D Scene Understanding

ACL 2023long

Visual spatial description (VSD) aims to generate texts that describe the spatial relations of the given objects within images. Existing VSD work merely models the 2D geometrical vision features, thus inevitably falling prey to the problem of skewed spatial understanding of target objects. In this w…

2023

Local-Global Defense against Unsupervised Adversarial Attacks on Graphs

AAAI 2023technical

Unsupervised pre-training algorithms for graph representation learning are vulnerable to adversarial attacks, such as first-order perturbations on graphs, which will have an impact on particular downstream applications. Designing an effective representation learning strategy against white-box attack…

Cited by 13SourcePDFScholar
2023

Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification

ICASSP 2023accepted

Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has diffi…

Cited by 0SourceScholar
2022

CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization

ICASSP 2022accepted

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-s…

Cited by 0SourceScholar
2022

DMANET: Deep Learning-Based Differential Microphone Arrays for Multi-Channel Speech Separation

ICASSP 2022accepted

In this paper, we develop a novel differential microphone arrays network (DMANet) for solving the multi-channel speech separation problem. In DMANet we explore a neural network combined to differential microphone arrays (DMAs) beamforming technique. Specifically, a sequence of differential operation…

Cited by 0SourceScholar
2022

Double Noise Mean Teacher Self-Ensembling Model for Semi-Supervised Tumor Segmentation

ICASSP 2022accepted

Accurate tumor segmentation of tumor images can assist doctors to diagnose diseases. However, achieving very high precision in tumor segmentation requires a large amount of annotated data, which is not easy for medical image data. In this paper, we present a novel double noise mean teacher self-ense…

Cited by 0SourceScholar
2022

Joint and Adversarial Training with ASR for Expressive Speech Synthesis

ICASSP 2022accepted

Style modeling is an important issue and has been proposed in expressive speech synthesis. In existing unsupervised methods, the style encoder extracts the latent representation from the reference audio as style information. However, the style information extracted from the style encoder will entang…

Cited by 0SourceScholar
2022

Visual Spatial Description: Controlled Spatial-Oriented Image-to-Text Generation

EMNLP 2022main

Image-to-text tasks such as open-ended image captioning and controllable image description have received extensive attention for decades. Here we advance this line of work further, presenting Visual Spatial Description (VSD), a new perspective for image-to-text toward spatial semantics. Given an ima…

2021

Portable Photoglottography for Monitoring Vocal Fold Vibrations in Speech Production

ICASSP 2021accepted

Photoglottography (PGG) is an effective method to monitor vocal fold vibrations via measuring light transmission across the glottis. The difficulty in operation however limits its wide use in speech studies. This paper is to realize a portable PGG (P-PGG) module with an audio interface to record glo…

Cited by 0SourceScholar
2021

Zero-Shot Voice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features

ICASSP 2021accepted

Zero-shot voice conversion (VC) where both source and target speakers are unseen in the training dataset has become a new research direction. Using speaker embeddings instead of one-hot vectors to represent speaker identity is a key point, which makes VC models work on unseen speakers. In our work,…

Cited by 0SourceScholar
2020

Retrieving Vocal-Tract Resonance and anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique

ICASSP 2020accepted

Vocal tract resonances give rise to core spectral information of speech signals. Linear prediction and cepstral methods are widely used for this purpose. However, both approaches are prone to fail as the fundamental frequency (F0) rises. In this study, a new cepstral method is developed combined wit…

Cited by 0SourceScholar
2020

Visual Encoding and Decoding of the Human Brain Based on Shared Features

IJCAI 2020poster

Using a convolutional neural network to build visual encoding and decoding models of the human brain is a good starting point for the study on relationship between deep learning and human visual cognitive mechanism. However, related studies have not fully considered their differences. In this paper,…

2019

Breast Cancer Detection Based on Merging Four Modes MRI Using Convolutional Neural Networks

ICASSP 2019accepted

The objective of the study is to develop a framework for automatic breast cancer detection with merging four imaging modes. Attempts were made for tumor classification and segmentation; using a multi-parametric Magnetic Resonance Imaging (MRI) method on breast tumors. MRI data of the breast were obt…

Cited by 0SourceScholar
2019

Glottographic and Aerodynamic Analysis on Consonant Aspiration and Onset F0 in Mandarin Chinese

ICASSP 2019accepted

Stop consonants in Mandarin Chinese are all voiceless at word-initial positions only showing aspirated and unaspirated distinctions. Between the two phonation types, voice onset time (VOT) shows a clear contrast in duration, whereas voice onset fundamental frequency (onset F0) does not, as seen in p…

Cited by 0SourceScholar
2016

Continuous ultrasound based tongue movement video synthesis from speech

ICASSP 2016accepted

The movement of tongue plays an important role in pronunciation. Visualizing the movement of tongue can improve speech intelligibility and also helps learning a second language. However, hardly any research has been investigated for this topic. In this paper, a framework to synthesize continuous ult…

Cited by 0SourceScholar
2015

Vocal responses to frequency modulated composite sinewaves via auditory and vibrotactile pathways

ICASSP 2015accepted

Feedback control mechanisms for speaking have been examined using the transformed auditory feedback (TAF) technique. Previous studies have shown that speakers demonstrate fundamental frequency (F0) changes when they monitor their voice with artificial alterations of F0. However, those studies undere…

Cited by 0SourceScholar