← Search

Li Lu

22 accepted papers

2026

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues

AAAI 2026technical

Toxic speech detection has become a crucial challenge in maintaining safe online communication environments. However, existing approaches to toxic speech detection often neglect the contribution of paralinguistic cues, such as emotion, intonation, and speech rate, which are key to detecting speech t

Cited by 0SourcePDFScholar
2026

HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection

ICML 2026poster

Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception. Although numerous audio deepfake detection (ADD) methods have been developed, most rely on local temporal/spectral features or pairwise relations, overlooking …

Cited by 0SourceScholar
2026

REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks

AAAI 2026technical

Large Language Models (LLMs) and causal learning each hold strong potential for clinical decision making (CDM). However, their synergy remains poorly understood, largely due to the lack of systematic benchmarks evaluating their integration in clinical risk prediction. In real-world healthcare, ident

Cited by 0SourcePDFScholar
2025

Integrating Personalized Spatio-Temporal Clustering for Next POI Recommendation

AAAI 2025technical

Location-Based Social Networks (LBSNs) offer a rich dataset of user activity at Points-of-Interest (POIs), making next POI recommendation a key task. Traditional algorithms face challenges due to broad searching scopes, affecting recommendation accuracy. Users tend to visit nearby POIs and show temp…

2025

SymmCompletion: High-Fidelity and High-Consistency Point Cloud Completion with Symmetry Guidance

AAAI 2025technical

Point cloud completion aims to recover a complete point shape from a partial point cloud. Although existing methods can form satisfactory point clouds in global completeness, they often lose the original geometry details and face the problem of geometric inconsistency between existing point clouds a…

2024

Exposing the Deception: Uncovering More Forgery Clues for Deepfake Detection

AAAI 2024technical

Deepfake technology has given rise to a spectrum of novel and compelling applications. Unfortunately, the widespread proliferation of high-fidelity fake videos has led to pervasive confusion and deception, shattering our faith that seeing is believing. One aspect that has been overlooked so far is t…

2023

Shift to Your Device: Data Augmentation for Device-Independent Speaker Verification Anti-Spoofing

ICASSP 2023accepted

This paper proposes a novel Deconvolution-enhanced data Augmentation method, DeAug, for ultrasonic-based speaker verification anti-spoofing systems to detect the liveness of voice sources in physical access, which aims to improve the performance of liveness detection on unseen devices where no data…

Cited by 0SourceScholar
2023

WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training

ICASSP 2023accepted

This paper aims to introduce a robust singing voice synthesis (SVS) system to produce very natural and realistic singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random area conditional discriminators to help supervise the acoust…

Cited by 0SourceScholar
2022

FBNet: Feedback Network for Point Cloud Completion

ECCV 2022poster

"The rapid development of point cloud learning has driven point cloud completion into a new era. However, the information flows of most existing completion methods are solely feedforward, and high-level information is rarely reused to improve low-level feature learning. To this end, we propose a nov…

2022

Online ECG Emotion Recognition for Unknown Subjects via Hypergraph-Based Transfer Learning

IJCAI 2022poster

Electrocardiogram (ECG) signal based cross-subject emotion recognition methods reduce the influence of individual differences using domain adaptation (DA) techniques. These methods generally assume that the entire unlabeled data of unknown target subjects are available in training phase. However, t…

Cited by 7SourcePDFScholar
2022

Zero-Shot Cross-Lingual Transfer Using Multi-Stream Encoder and Efficient Speaker Representation

ICASSP 2022accepted

We propose a novel method for zero-shot cross-lingual TTS task by using multi-stream text encoder and efficient speaker representation. Specifically, a unified multi-stream text encoder that takes both advantages of Transformer and CBHG is proposed to retain multiple hypotheses about input represent…

Cited by 0SourceScholar
2021

A New High Quality Trajectory Tiling Based Hybrid TTS In Real Time

ICASSP 2021accepted

A trajectory tiling based, hybrid TTS is revisited in this study for improving its synthesis performance. A combination of Transformer encoder and RNN based decoder architecture where two-level, at both word and Chinese phonetic alphabet letter levels, linguistic representation is exploited to gener…

Cited by 0SourceScholar
2021

Investigation of Fast and Efficient Methods for Multi-Speaker Modeling and Speaker Adaptation

ICASSP 2021accepted

In this paper, we propose a novel method for fast and efficient few-shot TTS task, which is able to disentangle linguistic and speaker representations. Specifically, an adversarial training strategy is firstly employed to wipe out speaker information from the linguistic representations. Then the spe…

Cited by 0SourceScholar
2020

An Improved Frame-Unit-Selection Based Voice Conversion System Without Parallel Training Data

ICASSP 2020accepted

A frame-unit-selection based voice conversion system proposed earlier by us is revisited here to enhance its performance in both speech naturalness and speaker similarity. Speaker independent, bilingual (Mandarin Chinese and American English) deep neural net (DNN) acoustic model’s output, frame-leve…

Cited by 1SourceScholar
2020

Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer

ICASSP 2020accepted

Although Transformer based neural end-to-end TTS model has demonstrated extreme effectiveness in capturing long-term dependencies and achieved state-of-the-art performance, it still suffers from two problems. 1) limited ability to model sequential and local structures in sequences; 2) heavily rely o…

Cited by 0SourceScholar
2019

Cross-culture Multimodal Emotion Recognition with Adversarial Learning

ICASSP 2019accepted

With the development of globalization, automatic emotion recognition has faced a new challenge in the multi-culture scenario - to generalize across different cultures. Previous works mainly rely on multi-cultural datasets to address the cross-culture discrepancy, which are expensive to collect. In t…

Cited by 0SourceScholar
2016

Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword search

ICASSP 2016accepted

In this paper, we propose a cross-lingual deep neural network (DNN) based submodular unbiased data selection approach for low-resource keyword search (KWS). A small amount (e.g. one hour) of transcribed data is used to conduct cross-lingual transfer. The frame-level senone sequence activated by the…

Cited by 0SourceScholar
2015

Submodular data selection with acoustic and phonetic features for automatic speech recognition

ICASSP 2015accepted

In this paper, we propose to use acoustic feature based submodular function optimization to select a subset of untranscribed data for manual transcription, and retrain the initial acoustic model with the additional transcribed data. The acoustic features are obtained from an unsupervised Gaussian mi…

Cited by 0SourceScholar