← Search

Qing He

33 accepted papers

2026

Scaling Speech Tokenizers with Diffusion Autoencoders

ICLR 2026poster

Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and (2) achieving low bit rates and low token rates. We propose Speech Diffusion To…

Cited by 2SourceScholar
2026

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

ICLR 2026poster

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal dynamics. Although large language models (LLMs) have shown promise i…

Cited by 0SourceScholar
2025

Controlling Large Language Models Through Concept Activation Vectors

AAAI 2025technical

As large language models (LLMs) are widely deployed across various domains, the ability to control their generated outputs has become more critical. This control involves aligning LLMs outputs with human values and ethical principles or customizing LLMs on specific topics or styles for individual us…

Cited by 1SourcePDFScholar
2025

Domain-aware Node Representation Learning for Graph Out-of-Distribution Generalization

ICASSP 2025accepted

Graph Neural Networks (GNNs) have demonstrated impressive success across diverse fields when data satisfies in-distribution (ID) assumption. Nevertheless, GNN performance significantly declines in cases of distribution shifts between training and testing graph data. This degradation primarily stems…

Cited by 0SourceScholar
2025

Dynamic Graph Learning with Static Relations for Credit Risk Assessment

AAAI 2025technical

Credit risk assessment has increasingly become a prominent research field due to the dramatically increased incidents of financial default. Traditional graph-based methods have been developed to detect defaulters within user-merchant commercial payment networks. However, these methods face challenge…

Cited by 0SourcePDFScholar
2025

Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation

ICASSP 2025accepted

Large language models (LLMs) have revolutionized natural language processing (NLP) with impressive performance across various text-based tasks. However, the extension of text-dominant LLMs to with speech generation tasks remains underexplored. In this work, we introduce a text-to-speech (TTS) system…

Cited by 6SourceScholar
2025

Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models

ACL 2025long

Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful technique for aligning large language models (LLMs) with human preferences. However, effectively aligning LLMs with diverse human preferences remains a significant challenge, particularly when they are conflict. To address t…

2024

Boosting the Adversarial Robustness of Graph Neural Networks: An OOD Perspective

ICLR 2024poster

Current defenses against graph attacks often rely on certain properties to eliminate structural perturbations by identifying adversarial edges from normal edges. However, this dependence makes defenses vulnerable to adaptive (white-box) attacks from adversaries with the same knowledge. Adversarial t…

2024

DRAMA: Dynamic Multi-Granularity Graph Estimate Retrieval over Tabular and Textual Question Answering

COLING 2024main

The TableTextQA task requires finding the answer to the question from a combination of tabular and textual data, which has been gaining increasing attention. The row-based approaches have demonstrated remarkable effectiveness. However, they suffer from the following limitations: (1) a lack of intera…

Cited by 2SourcePDFScholar
2024

EFSA: Towards Event-Level Financial Sentiment Analysis

ACL 2024long

In this paper, we extend financial sentiment analysis (FSA) to event-level since events usually serve as the subject of the sentiment in financial text. Though extracting events from the financial text may be conducive to accurate sentiment predictions, it has specialized challenges due to the lengt…

2024

F2GNN: An Adaptive Filter with Feature Segmentation for Graph-Based Fraud Detection

ICASSP 2024accepted

Graph Neural Networks (GNNs) have received remarkable success in identifying fraudulent activities on graphs. Most approaches leverage the full user feature together and aggregate the messages from its neighbors by a graph filter. However, due to the adversarial activities like the camouflage of fra…

Cited by 0SourceScholar
2024

Multi-Task Learning for Front-End Text Processing in TTS

ICASSP 2024accepted

We propose a multi-task learning (MTL) model for jointly performing three tasks that are commonly solved in a text-to-speech (TTS) front-end: text normalization (TN), part-of-speech (POS) tagging, and homograph disambiguation (HD). Our framework utilizes a tree-like structure with a trunk that learn…

Cited by 0SourceScholar
2024

Online Conversion Rate Prediction via Multi-Interval Screening and Synthesizing under Delayed Feedback

AAAI 2024technical

Due to the widespread adoption of the cost-per-action(CPA) display strategy that demands a real-time conversion rate prediction(CVR), delayed feedback is becoming one of the major challenges in online advertising. As the true labels of a significant quantity of samples are only available after long…

2024

Ultra-Lightweight Neural Differential DSP Vocoder for High Quality Speech Synthesis

ICASSP 2024accepted

Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing (DSP) based vocoder can be implemented via lightweight fast Fou…

Cited by 0SourceScholar
2023

Gradient-Adaptive Pareto Optimization for Constrained Reinforcement Learning

AAAI 2023technical

Constrained Reinforcement Learning (CRL) burgeons broad interest in recent years, which pursues maximizing long-term returns while constraining costs. Although CRL can be cast as a multi-objective optimization problem, it is still facing the key challenge that gradient-based Pareto optimization meth…

Cited by 6SourcePDFScholar
2023

Revisiting Graph Adversarial Attack and Defense From a Data Distribution Perspective

ICLR 2023poster

Recent studies have shown that structural perturbations are significantly effective in degrading the accuracy of Graph Neural Networks (GNNs) in the semi-supervised node classification (SSNC) task. However, why the gradient-based methods are so destructive is rarely explored. In this work, we discov…

Cited by 39SourcePDFScholar
2023

Self-Supervised Representations for Singing Voice Conversion

ICASSP 2023accepted

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Vec 2.0 have helped further the state-of-the-art. Though these methods produce mor…

Cited by 25SourceScholar
2023

Voice-Preserving Zero-Shot Multiple Accent Conversion

ICASSP 2023accepted

Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker’s accent. For native speakers, understanding or speaking a new accent is likewise a difficult task. An accent conversion system that changes a speaker’s accent b…

Cited by 26SourceScholar
2022

Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation Complexity

ICASSP 2022accepted

Low bitrate speech codecs have become an area of intense research. Traditional speech codecs, which use signal processing methods to encode and decode speech, often suffer from quality issues at low bitrates. A neural speech codec, which uses a deep neural network in the compression pipeline, can he…

Cited by 0SourceScholar
2022

Direct Speech-to-Speech Translation With Discrete Units

ACL 2022long

We present a direct speech-to-speech translation (S2ST) model that translates speech from one language to speech in another language without relying on intermediate text generation. We tackle the problem by first applying a self-supervised discrete speech encoder on the target speech and then traini…

2022

Mind the Gap: Cross-Lingual Information Retrieval with Hierarchical Knowledge Enhancement

AAAI 2022technical

Cross-Lingual Information Retrieval (CLIR) aims to rank the documents written in a language different from the user’s query. The intrinsic gap between different languages is an essential challenge for CLIR. In this paper, we introduce the multilingual knowledge graph (KG) to the CLIR task due to the…

Cited by 25SourcePDFScholar
2022

Multilingual Text-To-Speech Training Using Cross Language Voice Conversion And Self-Supervised Learning Of Speech Representations

ICASSP 2022accepted

State of the art text-to-speech (TTS) models can generate high fidelity monolingual speech, but it is still challenging to synthesize multilingual speech from the same speaker. One major hurdle is for training data. It’s hard to find speakers who have native proficiency in several languages. One way…

Cited by 0SourceScholar
2022

Vocbench: A Neural Vocoder Benchmark for Speech Synthesis

ICASSP 2022accepted

Neural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on synthesizing waveforms from low-dimensional representation, such as Mel-Spectrograms. In recent years, different approache…

Cited by 0SourceScholar
2021

AMA-GCN: Adaptive Multi-layer Aggregation Graph Convolutional Network for Disease Prediction

IJCAI 2021poster

Recently, Graph Convolutional Networks (GCNs) have proven to be a powerful mean for Computer Aided Diagnosis (CADx). This approach requires building a population graph to aggregate structural information, where the graph adjacency matrix represents the relationship between nodes. Until now, this adj…

Cited by 22SourcePDFScholar
2021

ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information

ACL 2021long

Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT, which incorporates both the glyph and pinyin information of…

2021

ConRPG: Paraphrase Generation using Contexts as Regularizer

EMNLP 2021main

A long-standing issue with paraphrase generation is the lack of reliable supervision signals. In this paper, we propose a new unsupervised paradigm for paraphrase generation based on the assumption that the probabilities of generating two sentences with the same meaning given the same context should…

Cited by 26SourcePDFScholar
2021

Discerning Decision-Making Process of Deep Neural Networks with Hierarchical Voting Transformation

NeurIPS 2021poster

Neural network based deep learning techniques have shown great success for numerous applications. While it is expected to understand their intrinsic decision-making processes, these deep neural networks often work in a black-box way. To this end, in this paper, we aim to discern the decision-making…

2021

Multi-Rate Attention Architecture for Fast Streamable Text-to-Speech Spectrum Modeling

ICASSP 2021accepted

Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates the actual audio. High-quality spectrum models usually incorporate the encoder-decoder architecture with self-attention…

Cited by 0SourceScholar
2021

PENS: A Dataset and Generic Framework for Personalized News Headline Generation

ACL 2021long

In this paper, we formulate the personalized news headline generation problem whose goal is to output a user-specific title based on both a user’s reading interests and a candidate news body to be exposed to her. To build up a benchmark for this problem, we publicize a large-scale dataset named PENS…

2021

Self Question-answering: Aspect-based Sentiment Analysis by Role Flipped Machine Reading Comprehension

EMNLP 2021finding

The pivot for the unified Aspect-based Sentiment Analysis (ABSA) is to couple aspect terms with their corresponding opinion terms, which might further derive easier sentiment predictions. In this paper, we investigate the unified ABSA task from the perspective of Machine Reading Comprehension (MRC)…

Cited by 20SourcePDFScholar
2020

Trust the Model When It Is Confident: Masked Model-based Actor-Critic

NeurIPS 2020poster

It is a popular belief that model-based Reinforcement Learning (RL) is more sample efficient than model-free RL, but in practice, it is not always true due to overweighed model errors. In complex and noisy settings, model-based RL tends to have trouble using the model if it does not know when to tru…

Cited by 61SourcePDFScholar