← Search

Ya Li

24 accepted papers

2026

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

ICML 2026poster

Spoken Language Models (SLMs) revolutionize speech synthesis by bypassing traditional linguistic front-ends, yet they remain limited by the digital resource disparities across languages. We investigate these challenges within the Southeast Asian linguistic landscape, using the phonetically complex T…

Cited by 0SourceScholar
2026

FAKE SPEECH WILD: DETECTING DEEPFAKE SPEECH ON SOCIAL MEDIA PLATFORM

ICASSP 2026poster

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios.…

Cited by 0SourcePDFScholar
2026

ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs

ICLR 2026poster

Large Language Models (LLMs) excel in various domains, but their safe deployment faces the challenge of balancing safety and utility. Existing alignment strategies often strengthen refusal mechanisms to reduce harmful outputs, but harmless instructions with superficial risky words are mistakenly rej…

Cited by 0SourcecodeScholar
2025

Beyond Surface Simplicity: Revealing Hidden Reasoning Attributes for Precise Commonsense Diagnosis

ACL 2025long

Commonsense question answering (QA) are widely used to evaluate the commonsense abilities of large language models. However, answering commonsense questions correctly requires not only knowledge but also reasoning—even for seemingly simple questions. We demonstrate that such hidden reasoning attribu…

Cited by 0SourcePDFScholar
2025

Controllable 3D Dance Generation Using Diffusion-Based Transformer U-Net

AAAI 2025technical

Recently, dance generation has attracted increasing interest. In particular, the success of diffusion models in image generation has led to the emergence of dance generation systems based on the diffusion framework. However, these systems lack controllability, which limits their practical applicatio…

Cited by 0SourcePDFScholar
2025

DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech

ICASSP 2025accepted

Traditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often…

Cited by 0SourceScholar
2025

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

ICML 2025poster

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-apprai…

2024

Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis

ICASSP 2024accepted

Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into enhancing context comprehension, context representation still lacks effective r…

Cited by 0SourceScholar
2024

Frame-Level Emotional State Alignment Method for Speech Emotion Recognition

ICASSP 2024accepted

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audio have affective states consistent with utterance-level label, which makes it di…

Cited by 0SourceScholar
2023

M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis

ICASSP 2023accepted

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and om…

Cited by 0SourceScholar
2022

Automatic Depression Level Assessment from Speech By Long-Term Global Information Embedding

ICASSP 2022accepted

Depression is a serious mood disorder which brings negative effects on people's social activities. Therefore, growing attention has been paid to automatic depression assessment, especially from speech. However, most of the previous work uses hand-crafted features or deep neural network-based feature…

Cited by 0SourceScholar
2022

Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional Network

ICASSP 2022accepted

Automated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classificati…

Cited by 0SourceScholar
2022

Towards Lightweight Black-Box Attack Against Deep Neural Networks

NeurIPS 2022accept

Black-box attacks can generate adversarial examples without accessing the parameters of target model, largely exacerbating the threats of deployed deep neural networks (DNNs). However, previous works state that black-box attacks fail to mislead target models when their training data and outputs are…

Cited by 23SourcePDFScholar
2021

Cross Attention Augmented Transducer Networks for Simultaneous Translation

EMNLP 2021main

This paper proposes a novel architecture, Cross Attention Augmented Transducer (CAAT), for simultaneous translation. The framework aims to jointly optimize the policy and translation models. To effectively consider all possible READ-WRITE simultaneous translation action paths, we adapt the online au…

2020

Dual-Path Distillation: A Unified Framework to Improve Black-Box Attacks

ICML 2020poster

We study the problem of constructing black-box adversarial attacks, where no model information is revealed except for the feedback knowledge of the given inputs. To obtain sufficient knowledge for crafting adversarial examples, previous methods query the target model with inputs that are perturbed w…

Cited by 17SourcePDFScholar
2020

Transferable, Controllable, and Inconspicuous Adversarial Attacks on Person Re-identification With Deep Mis-Ranking

CVPR 2020oral

The success of DNNs has driven the extensive applications of person re-identification (ReID) into a new era. However, whether ReID inherits the vulnerability of DNNs remains unexplored. To examine the robustness of ReID systems is rather important because the insecurity of ReID systems may cause sev…

Cited by 106PDFcodeScholar
2019

Discriminative Video Representation with Temporal Order for Micro-expression Recognition

ICASSP 2019accepted

Micro-expression recognition is a challenging task due to its low intensity and short duration and how to extract the subtle facial changes is a key issue in this field. Although there are many methods attempt to cope with this problem, they are difficult to encode the temporal order of all frames i…

Cited by 0SourceScholar
2018

Deep Domain Generalization via Conditional Invariant Adversarial Networks

ECCV 2018poster

Domain generalization aims to learn a classification model from multiple source domains and generalize it to unseen target domains. A critical problem in domain generalization involves learning domain-invariant representations. Let $X$ and $Y$ denote the features and the labels, respectively. Under…

Cited by 890SourcePDFScholar
2018

End-to-End Continuous Emotion Recognition from Video Using 3D Convlstm Networks

ICASSP 2018accepted

Conventional continuous emotion recognition consists of feature extraction step followed by regression step. However, the objective of the two steps is not consistent as they are parted. Besides, there is still no consensus about appropriate emotional features. In this study, we propose an end-to-en…

Cited by 0SourceScholar
2016

Long short term memory recurrent neural network based encoding method for emotion recognition in video

ICASSP 2016accepted

Human emotion is a temporally dynamic event which can be inferred from both audio and video feature sequences. In this paper we investigate the long short term memory recurrent neural network (LSTM-RNN) based encoding method for category emotion recognition in the video. LSTM-RNN is able to incorpor…

Cited by 0SourceScholar