← Search

Jing Xiao

119 accepted papers

2026

A Progressive Evidence Localization Framework Based on Wasserstein Gradient Flows for Document Visual Question Answering

ICML 2026poster

Precise evidence region localization in Document Visual Question Answering (DocVQA) is crucial for improving model interpretability and reliability. However, most existing approaches rely on single-step localization, which struggles to effectively distinguish true evidence from irrelevant content wh…

Cited by 0SourceScholar
2026

CHARM: Collaborative Harmonization Across Arbitrary Modalities for Modality-Agnostic Semantic Segmentation

AAAI 2026technical

Modality-agnostic Semantic Segmentation (MaSS) aims to achieve robust scene understanding across arbitrary combinations of input modality. Existing methods typically rely on explicit feature alignment to achieve modal homogenization, which dilutes the distinctive strengths of each modality and destr

Cited by 0SourcePDFScholar
2026

Learning to Generate Structured Meshes with In-Context: Toward Generalization in Mesh Generation

AAAI 2026technical

Structured mesh generation serves as a crucial preprocessing step in numerical simulations and can be formulated as a mapping problem from geometry to structured mesh. Existing approaches typically establish an isolated mapping for each geometry. This geometry-specific paradigm fails to capture and

Cited by 0SourcePDFScholar
2026

Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry

ICLR 2026poster

Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this “LLM-as-a-Judge” paradigm is costly, opaque, and sensitive to prompt design. In this work, we investigate whether smaller models can serve as efficient evaluators by leveraging internal representations…

Cited by 0SourcecodeScholar
2025

ACCon: Angle-Compensated Contrastive Regularizer for Deep Regression

AAAI 2025technical

In deep regression, capturing the relationship among continuous labels in feature space is a fundamental challenge that has attracted increasing interest. Addressing this issue can prevent models from converging to suboptimal solutions across various regression tasks, leading to improved performance…

Cited by 0SourcePDFScholar
2025

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

ACL 2025long

Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and human-like responses, their **lack of controllability** remains a key challenge, often leading to unfocused conversations or task failure. To address this, we…

2025

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

CVPR 2025poster

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS keypoint movements,…

2025

Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models

ACL 2025finding

Large language models (LLMs) often exhibit Context Faithfulness Hallucinations, where outputs deviate from retrieved information due to incomplete context integration. Our analysis reveals a strong correlation between token-level uncertainty and hallucinations. We hypothesize that attention mechanis…

Cited by 0SourcePDFScholar
2025

EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed

ICASSP 2025accepted

Non-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. In this paper, we propose a single-step NAR A…

Cited by 0SourceScholar
2025

GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression

EMNLP 2025

Recent studies have demonstrated that many layers are functionally redundant in large language models (LLMs), enabling model compression by removing these layers to reduce inference cost. While such approaches can improve efficiency, indiscriminate layer pruning often results in significant performa

2025

Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC Loss

ICASSP 2025accepted

Contextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the…

Cited by 0SourceScholar
2025

LEF-TTS: Lightweight and Efficient End-to-End Text-to-Speech Synthesis With Multi-Stream Generator

ICASSP 2025accepted

Recently, the field of Text-to-speech synthesis has been predominantly characterized by end-to-end models, with the quality of speech generated by these models becoming increasingly comparable to that of human speech. In this work, we propose a Lightweight and Efficient Text-to-speech model, a fast…

Cited by 0SourceScholar
2025

Open-world Radio Frequency Fingerprint Identification via Augmented Semi-supervised Learning

AAAI 2025technical

In complex electromagnetic environments, the identification and differentiation of diverse radio frequency (RF) emitters become particularly crucial. Existing RF fingerprinting methods demonstrate limitations when dealing with numerous unknown emitters, making it challenging for accurate classificat…

2025

RUNA: Object-Level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations

AAAI 2025technical

Enabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do not receive supervisory signals from unfamiliar data, leading to overly confident predictions regarding OOD objects. Despi…

Cited by 0SourcePDFScholar
2025

SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation

ICASSP 2025accepted

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. While audio-driven talking face generation has seen notable advancements, existing m…

Cited by 0SourceScholar
2025

Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation

ICASSP 2025accepted

The rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation…

Cited by 0SourceScholar
2025

Token-Level Contextual Network with Ladder-Shaped Attention for End-to-End ASR

ICASSP 2025accepted

Contextual automatic speech recognition (ASR) plays an increasingly important role in addressing the long-tail issues of general ASR. In the past, contextual ASR mainly focused on phrase-level discussions, providing a convenient way to handle biasing phrases. This paper introduces a new contextual n…

Cited by 0SourceScholar
2024

Bidirectional Autoregessive Diffusion Model for Dance Generation

CVPR 2024poster

Dance serves as a powerful medium for expressing human emotions but the lifelike generation of dance is still a considerable challenge. Recently diffusion models have showcased remarkable generative abilities across various domains. They hold promise for human motion generation due to their adaptabl…

Cited by 8SourcePDFScholar
2024

Boosting Image Quality Assessment through Efficient Transformer Adaptation with Local Feature Enhancement

CVPR 2024poster

Image Quality Assessment (IQA) constitutes a fundamental task within the field of computer vision yet it remains an unresolved challenge owing to the intricate distortion conditions diverse image contents and limited availability of data. Recently the community has witnessed the emergence of numerou…

2024

Co-speech Gesture Video Generation with 3D Human Meshes

ECCV 2024poster

"Co-speech gesture video generation is an enabling technique for many digital human applications. Substantial progress has been made in creating high-quality talking head videos. However, existing hand gesture video generation methods are primarily limited by the widely adopted 2D skeleton-based ges…

Cited by 1SourcePDFScholar
2024

DFlow: A Generative Model Combining Denoising AutoEncoder and Normalizing Flow for High Fidelity Waveform Generation

ICML 2024poster

In this work, we present DFlow, a novel generative framework that combines Normalizing Flow (NF) with a Denoising AutoEncoder (DAE), for high-fidelity waveform generation. With a tactfully designed structure, DFlow seamlessly integrates the capabilities of both NF and DAE, resulting in a significant…

Cited by 0SourcePDFScholar
2024

ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

ICASSP 2024accepted

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech synthesis model that leverages Speech Emotion Diarization (…

Cited by 0SourceScholar
2024

ESVC: Combining Adaptive Style Fusion and Multi-Level Feature Disentanglement for Expressive Singing Voice Conversion

ICASSP 2024accepted

Nowadays, singing voice conversion (SVC) has made great strides in both naturalness and similarity for common SVC with a neutral expression. However, besides singer identity, emotional expression is also essential to convey the singer’s emotions and attitudes, but current SVC systems can not effecti…

Cited by 0SourceScholar
2024

EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model

ICASSP 2024accepted

In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with…

Cited by 0SourceScholar
2024

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

NAACL 2024long

In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curat…

2024

IDEAW: Robust Neural Audio Watermarking with Invertible Dual-Embedding

EMNLP 2024main

The audio watermarking technique embeds messages into audio and accurately extracts messages from the watermarked audio. Traditional methods develop algorithms based on expert experience to embed watermarks into the time-domain or transform-domain of signals. With the development of deep neural netw…

2024

INCPrompt: Task-Aware Incremental Prompting for Rehearsal-Free Class-Incremental Learning

ICASSP 2024accepted

This paper introduces INCPrompt, an innovative continual learning solution that effectively addresses catastrophic forgetting. INCPrompt’s key innovation lies in its use of adaptive key-learner and task-aware prompts that capture task-relevant information. This unique combination encapsulates genera…

Cited by 0SourceScholar
2024

Improving Attention-Based End-to-End Speech Recognition by Monotonic Alignment Attention Matrix Reconstruction

ICASSP 2024accepted

In automatic speech recognition (ASR) task, the output sequence should correspond to a linear transcription of the input sequence. Lots of works have been done to learn the monotonic alignment in end-to-end (E2E) ASR model, but their methods mainly focus on streaming propose and usually result in a…

Cited by 0SourceScholar
2024

Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval

ICASSP 2024accepted

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides,…

Cited by 0SourceScholar
2024

Leveraging Biases in Large Language Models: "bias-kNN" for Effective Few-Shot Learning

ICASSP 2024accepted

Large Language Models (LLMs) have shown significant promise in various applications, including zero-shot and few-shot learning. However, their performance can be hampered by inherent biases. Instead of traditionally sought methods that aim to minimize or correct these biases, this study introduces a…

Cited by 0SourceScholar
2024

P2DT: Mitigating Forgetting in Task-Incremental Learning with Progressive Prompt Decision Transformer

ICASSP 2024accepted

Catastrophic forgetting poses a substantial challenge for managing intelligent agents controlled by a large model, causing performance degradation when these agents face new tasks. In our work, we propose a novel solution - the Progressive Prompt Decision Transformer (P2DT). This method enhances a t…

Cited by 0SourceScholar
2024

Prior Relational Schema Assists Effective Contrastive Learning for Inductive Knowledge Graph Completion

COLING 2024main

Knowledge Graph Completion (KGC) is a task aimed at uncovering the inherent relationships among known knowledge triplets in a Knowledge Graph (KG) and subsequently predicting missing links. Presently, there is a rising interest in inductive knowledge graph completion, where missing links may pertain…

2023

Assessor360: Multi-sequence Network for Blind Omnidirectional Image Quality Assessment

NeurIPS 2023poster

Blind Omnidirectional Image Quality Assessment (BOIQA) aims to objectively assess the human perceptual quality of omnidirectional images (ODIs) without relying on pristine-quality image information. It is becoming more significant with the increasing advancement of virtual reality (VR) technology. H…

2023

Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation

ICASSP 2023accepted

While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to ha…

Cited by 0SourceScholar
2023

Detecting Out-of-Distribution Examples Via Class-Conditional Impressions Reappearing

ICASSP 2023accepted

Out-of-distribution (OOD) detection aims at enhancing standard deep neural networks to distinguish anomalous inputs from original training data. Previous progress has introduced various approaches where the in-distribution training data and even several OOD examples are prerequisites. However, due t…

Cited by 0SourceScholar
2023

Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross Entropy

ICASSP 2023accepted

Because of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autoregressive models. In this work, we present dynamic alignment Mask CTC, introducing two methods: (1) Aligned Cross Entrop…

Cited by 0SourceScholar
2023

Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response Retrieval

ICASSP 2023accepted

Deep neural networks have achieved remarkable performance in retrieval-based dialogue systems, but they are shown to be ill calibrated. Though basic calibration methods like Monte Carlo Dropout and Ensemble can calibrate well, these methods are time-consuming in the training or inference stages. To…

Cited by 0SourceScholar
2023

Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification

ICASSP 2023accepted

Data-Free Knowledge Distillation (DFKD) has recently attracted growing attention in the academic community, especially with major breakthroughs in computer vision. Despite promising results, the technique has not been well applied to audio and signal processing. Due to the variable duration of audio…

Cited by 0SourceScholar
2023

FedET: A Communication-Efficient Federated Class-Incremental Learning Framework Based on Enhanced Transformer

IJCAI 2023poster

Federated Learning (FL) has been widely concerned for it enables decentralized learning while ensuring data privacy. However, most existing methods unrealistically assume that the classes encountered by local clients are fixed over time. After learning new classes, this impractical assumption will m…

Cited by 30SourcePDFScholar
2023

GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution Detection

NeurIPS 2023poster

Detecting out-of-distribution (OOD) examples is crucial to guarantee the reliability and safety of deep neural networks in real-world settings. In this paper, we offer an innovative perspective on quantifying the disparities between in-distribution (ID) and OOD data---analyzing the uncertainty that…

2023

Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial Representations

ICASSP 2023accepted

Using deep learning methods to classify EEG signals can accurately identify people’s emotions. However, existing studies have rarely considered the application of the information in another domain’s representations to feature selection in the time-frequency domain. We propose a classification networ…

Cited by 0SourceScholar
2023

Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations Perspective

ICASSP 2023accepted

Music genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use audio content and lyrics content inefficiently. In addition, as…

Cited by 0SourceScholar
2023

Learning Speech Representations with Flexible Hidden Feature Dimensions

ICASSP 2023accepted

Non-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance…

Cited by 0SourceScholar
2023

On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval Models

AAAI 2023technical

Deep neural retrieval models have amply demonstrated their power but estimating the reliability of their predictions remains challenging. Most dialog response retrieval models output a single score for a response on how relevant it is to a given question. However, the bad calibration of deep neural…

2023

Only a Few Classes Confusing: Pixel-Wise Candidate Labels Disambiguation for Foggy Scene Understanding

AAAI 2023technical

Not all semantics become confusing when deploying a semantic segmentation model for real-world scene understanding of adverse weather. The true semantics of most pixels have a high likelihood of appearing in the few top classes according to confidence ranking. In this paper, we replace the one-hot p…

Cited by 9SourcePDFScholar
2023

PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual Adapter

EMNLP 2023long main

The Retrieval Question Answering (ReQA) task employs the retrieval-augmented framework, composed of a retriever and generator. The generators formulate the answer based on the documents retrieved by the retriever. Incorporating Large Language Models (LLMs) as generators is beneficial due to their ad…

Cited by 0SourceScholar
2023

Predicting Center of Mass by Iterative Pushing for Object Transportation and Manipulation

IROS 2023poster

Robotic manipulation tasks rely on a plethora of environmental and payload information. One critical piece of information for accurate manipulation is the center of mass (CoM) of the object, which is essential for estimating the dynamic response of the system and determining the payload placement. T…

Cited by 3SourceScholar
2023

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

ICASSP 2023accepted

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to further deliver the speaker’s questioning intention while tran…

Cited by 0SourceScholar
2023

VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector Quantization

ICASSP 2023accepted

Voice Conversion(VC) refers to converting the voice characteristics of audio to another one as it is said by other people. Recently, more and more studies have focused on disentangle-based VC, which separates the timbre and linguistic content information from an audio signal to effectively achieve V…

Cited by 0SourceScholar
2022

An Augmented Benchmark Dataset for Geometric Question Answering through Dual Parallel Text Encoding

COLING 2022main

Automatic math problem solving has attracted much attention of NLP researchers recently. However, most of the works focus on the solving of Math Word Problems (MWPs). In this paper, we study on the Geometric Problem Solving based on neural networks. Solving geometric problems requires the integratio…

2022

Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning

ICASSP 2022accepted

Voice Conversion(VC) refers to changing the timbre of a speech while retaining the discourse content. Recently, many works have focused on disentangle-based learning techniques to separate the timbre and the linguistic content information from a speech signal. Once successful, voice conversion will…

Cited by 0SourceScholar
2022

DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised Learning

ICASSP 2022accepted

Any-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech consists of content and speaker style information and aims to unta…

Cited by 0SourceScholar
2022

ElasticMVS: Learning elastic part representation for self-supervised multi-view stereopsis

NeurIPS 2022accept

Self-supervised multi-view stereopsis (MVS) attracts increasing attention for learning dense surface predictions from only a set of images without onerous ground-truth 3D training data for supervision. However, existing methods highly rely on the local photometric consistency, which fails to identif…

Cited by 9SourcePDFScholar
2022

Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis

EMNLP 2022finding

Multimodal speech emotion recognition (SER) and sentiment analysis (SA) are important techniques for human-computer interaction. Most existing multimodal approaches utilize either shallow cross-modal fusion of pretrained features, or deep cross-modal fusion with raw features. Recently, attempts have…

2022

Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual Learning

CVPR 2022poster

Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal…

Cited by 44PDFcodeScholar
2022

nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speech

ICASSP 2022accepted

Multi-speaker text-to-speech (TTS) using a few adaption data is a challenge in practical applications. To address that, we propose a zero-shot multi-speaker TTS, named nnSpeech, that could synthesis a new speaker voice without fine-tuning and using only one adaption utterance. Compared with using a…

Cited by 0SourceScholar
2022

r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled Noise Introducing and Contextual Information Incorporation

ICASSP 2022accepted

Grapheme-to-phoneme (G2P) conversion is the process of converting the written form of words to their pronunciations. It has an important role for text-to-speech (TTS) synthesis and automatic speech recognition (ASR) systems. In this paper, we aim to evaluate and enhance the robustness of G2P models.…

Cited by 0SourceScholar
2021

3D Graph Anatomy Geometry-Integrated Network for Pancreatic Mass Segmentation, Diagnosis, and Quantitative Patient Management

CVPR 2021poster

The pancreatic disease taxonomy includes ten types of masses (tumors or cysts) [20, 8]. Previous work focuses on developing segmentation or classification methods only for certain mass types. Differential diagnosis of all mass types is clinically highly desirable [20] but has not been investigated u…

Cited by 47PDFScholar
2021

A Neural Transition-based Joint Model for Disease Named Entity Recognition and Normalization

ACL 2021long

Disease is one of the fundamental entities in biomedical research. Recognizing such entities from biomedical text and then normalizing them to a standardized disease vocabulary offer a tremendous opportunity for many downstream applications. Previous studies have demonstrated that joint modeling of…

Cited by 19SourcePDFScholar
2021

An Alignment-Agnostic Model for Chinese Text Error Correction

EMNLP 2021finding

This paper investigates how to correct Chinese text errors with types of mistaken, missing and redundant characters, which are common for Chinese native speakers. Most existing models based on detect-correct framework can correct mistaken characters, but cannot handle missing or redundant characters…

Cited by 5SourcePDFScholar
2021

Automatic Vertebra Localization and Identification in CT by Spine Rectification and Anatomically-Constrained Optimization

CVPR 2021poster

Accurate vertebra localization and identification are required in many clinical applications of spine disorder diagnosis and surgery planning. However, significant challenges are posed in this task by highly varying pathologies (such as vertebral compression fracture, scoliosis, and vertebral fixati…

Cited by 35PDFcodeScholar
2021

CASS-NAT: CTC Alignment-Based Single Step Non-Autoregressive Transformer for Speech Recognition

ICASSP 2021accepted

We propose a CTC alignment-based single step non-autoregressive transformer (CASS-NAT) for speech recognition. Specifically, the CTC alignment contains the information of (a) the number of tokens for decoder input, and (b) the time span of acoustics for each token. The information are used to extrac…

Cited by 0SourceScholar
2021

Deep Lesion Tracker: Monitoring Lesions in 4D Longitudinal Imaging Studies

CVPR 2021poster

Monitoring treatment response in longitudinal studies plays an important role in clinical practice. Accurately identifying lesions across serial imaging follow-up is the core to the monitoring procedure. Typically this incorporates both image and anatomical considerations. However, matching lesions…

Cited by 48PDFcodeScholar
2021

Efficient Client Contribution Evaluation for Horizontal Federated Learning

ICASSP 2021accepted

In federated learning (FL), fair and accurate measurement of the contribution of each federated participant is of great significance. The level of contribution not only provides a rational metric for distributing financial benefits among federated participants, but also helps to discover malicious p…

Cited by 0SourceScholar
2021

EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture

ICML 2021spotlight

In this work, we address the Text-to-Speech (TTS) task by proposing a non-autoregressive architecture called EfficientTTS. Unlike the dominant non-autoregressive TTS models, which are trained with the need of external aligners, EfficientTTS optimizes all its parameters with a stable, end-to-end trai…

2021

Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation

ICASSP 2021accepted

Knowledge distillation refers to a technique of transferring the knowledge from a large learned model or an ensemble of learned models to a small model. This method relies on access to the original training set, which might not always be available. A possible solution is a data-free adversarial dist…

Cited by 0SourceScholar
2021

Enhancing Dual-Encoders with Question and Answer Cross-Embeddings for Answer Retrieval

EMNLP 2021finding

Dual-Encoders is a promising mechanism for answer retrieval in question answering (QA) systems. Currently most conventional Dual-Encoders learn the semantic representations of questions and answers merely through matching score. Researchers proposed to introduce the QA interaction features in scorin…

2021

Image Inpainting Guided by Coherence Priors of Semantics and Textures

CVPR 2021poster

Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we…

Cited by 120PDFScholar
2021

Improving Neural Text Normalization with Partial Parameter Generator and Pointer-Generator Network

ICASSP 2021accepted

Text Normalization (TN) is an essential part in conversational systems like text-to-speech synthesis (TTS) and automatic speech recognition (ASR). It is a process of transforming non-standard words (NSW) into a representation of how the words are to be spoken. Existing approaches to TN are mainly ru…

Cited by 0SourceScholar
2021

Joint Intent Detection and Slot Filling Based on Continual Learning Model

ICASSP 2021accepted

Slot filling and intent detection have become a significant theme in the field of natural language understanding. Even though slot filling is intensively associated with intent detection, the characteristics of the information required for both tasks are different while most of those approaches may…

Cited by 0SourceScholar
2021

LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation

ICASSP 2021accepted

In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use of unified convolution kernels in WaveNet to capture the dependencies of arbitrary waveform, the location-variable convol…

Cited by 0SourceScholar
2021

Leveraging Large-Scale Weakly Labeled Data for Semi-Supervised Mass Detection in Mammograms

CVPR 2021poster

Mammographic mass detection is an integral part of a computer-aided diagnosis system. Annotating a large number of mammograms at pixel-level in order to train a mass detection model in a fully supervised fashion is costly and time-consuming. This paper presents a novel self-training framework for se…

Cited by 14PDFScholar
2021

Multi-Grained Knowledge Distillation for Named Entity Recognition

NAACL 2021long

Although pre-trained big models (e.g., BERT, ERNIE, XLNet, GPT3 etc.) have delivered top performance in Seq2seq modeling, their deployments in real-world applications are often hindered by the excessive computations and memory demand involved. For many applications, including named entity recognitio…

Cited by 17SourcePDFScholar
2021

Network Pruning Using Linear Dependency Analysis on Feature Maps

ICASSP 2021accepted

Network pruning can be achieved by removing redundant channels. In this paper, we regard a channel ‘redundant’ if its output is linearly dependent with respect to those of other channels. Inspired by this, we propose an efficient pruning method, named as LDFM, by linear dependency analysis on all th…

Cited by 0SourceScholar
2021

PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling Check

ACL 2021long

Chinese Spelling Check (CSC) is a challenging task due to the complex characteristics of Chinese characters. Statistics reveal that most Chinese spelling errors belong to phonological or visual errors. However, previous methods rarely utilize phonological and morphological knowledge of Chinese chara…

2021

SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech Recognition

ICASSP 2021accepted

Inspired by the contrastive predictive coding (CPC), we propose a feature representation scheme for automatic speech recognition (ASR), which encodes sequential dependency information from raw audio signals. Following the original CPC, for a given frame, mutual information (MI) lower bound is maximi…

Cited by 0SourceScholar
2021

Understanding Gradient Clipping In Incremental Gradient Methods

AISTATS 2021poster

We provide a theoretical analysis on how gradient clipping affects the convergence of the incremental gradient methods on minimizing an objective function that is the sum of a large number of component functions. We show that clipping on gradients of component functions leads to bias on the descent…

Cited by 49SourcePDFScholar
2021

Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition

ICASSP 2021accepted

Self-attention models have been successfully applied in end-to-end speech recognition systems, which greatly improve the performance of recognition accuracy. However, such attention-based models cannot be used in online speech recognition, because these models usually have to utilize a whole acousti…

Cited by 0SourceScholar
2021

Unsupervised Learning for Multi-Style Speech Synthesis with Limited Data

ICASSP 2021accepted

Existing multi-style speech synthesis methods require either style labels or large amounts of unlabeled training data, making data acquisition difficult. In this paper, we present an unsupervised multi-style speech synthesis method that can be trained with limited data. We leverage instance discrimi…

Cited by 0SourceScholar
2021

Window Loss for Bone Fracture Detection and Localization in X-ray Images with Point-based Annotation

AAAI 2021technical

Object detection methods are widely adopted for computer-aided diagnosis using medical images. Anomalous findings are usually treated as objects that are described by bounding boxes. Yet, many pathological findings, e.g., bone fractures, cannot be clearly defined by bounding boxes, owing to consider…

Cited by 2SourcePDFScholar
2020

A Robust Speaker Clustering Method Based on Discrete Tied Variational Autoencoder

ICASSP 2020accepted

Recently, the speaker clustering model based on aggregation hierarchy cluster (AHC) is a common method to solve two main problems: no preset category number clustering and fix category number clustering. In general, model takes features like i-vectors as input of probability and linear discriminant…

Cited by 0SourceScholar
2020

Aligntts: Efficient Feed-Forward Text-to-Speech System Without Explicit Alignment

ICASSP 2020accepted

Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-spectrum from a sequence of characters, and the duration of each character is determined by a duration predictor. Instea…

Cited by 0SourceScholar
2020

Anatomy-Aware Siamese Network: Exploiting Semantic Asymmetry for Accurate Pelvic Fracture Detection in X-ray Images

ECCV 2020poster

Trauma PXR are essential for instantaneous pelvic bone fracture detection. However, small, pathologically critical fractures can be missed, even by experienced clinicians, under the very limited diagnosis times allowed in urgent care. As a result, fracture CAD has very high demands to save time and…

Cited by 43SourcePDFScholar
2020

Co-Heterogeneous and Adaptive Segmentation from Multi-Source and Multi-Phase CT Imaging Data: A Study on Pathological Liver and Lesion Segmentation

ECCV 2020poster

Within medical imaging, organ/pathology segmentation models trained on current publicly available and fully-annotated datasets usually do not well-represent the heterogeneous modalities, phases, pathologies, and clinical scenarios encountered in real environments. On the other hand, there are tremen…

Cited by 36SourcePDFScholar
2020

Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on Flow

ICASSP 2020accepted

In this work, we propose Flow-TTS, a non-autoregressive end-to-end neural TTS model based on generative flow. Unlike other non-autoregressive models, Flow-TTS can achieve high-quality speech generation by using a single feed-forward network. To our knowledge, Flow-TTS is the first TTS model utilizin…

Cited by 0SourceScholar
2020

Generating Reasonable Legal Text through the Combination of Language Modeling and Question Answering

IJCAI 2020poster

Due to the improvement of Language Modeling, the emerging NLP assistant tools aiming for text generation greatly reduce the human workload on writing documents. However, the generation of legal text faces greater challenges than ordinary texts because of its high requirement for keeping logic reason…

2020

GraphTTS: Graph-to-Sequence Modelling in Neural Text-to-Speech

ICASSP 2020accepted

This paper leverages the graph-to-sequence method in neural text-to-speech (GraphTTS), which maps the graph embedding of the input sequence to spectrograms. The graphical inputs consist of node and edge representations constructed from input texts. The encoding of these graphical inputs incorporates…

Cited by 0SourceScholar
2020

Guidance and Evaluation: Semantic-Aware Image Inpainting for Mixed Scenes

ECCV 2020poster

Completing a corrupted image with correct structures and reasonable textures for a mixed scene remains an elusive challenge. Since the missing hole in a mixed scene of a corrupted image often contains various semantic information, conventional two-stage approaches utilizing structural information of…

Cited by 154SourcePDFScholar
2020

JSSR: A Joint Synthesis, Segmentation, and Registration System for 3D Multi-Modal Image Alignment of Large-scale Pathological CT Scans

ECCV 2020poster

Segmentation, and Registration System for 3D Multi-Modal Image Alignment of Large-scale Pathological CT Scans","Multi-modal image registration is a challenging problem that is also an important clinical task for many real applications and scenarios. As a first step in analysis, deformable registrati…

Cited by 30SourcePDFScholar
2020

Organ at Risk Segmentation for Head and Neck Cancer Using Stratified Learning and Neural Architecture Search

CVPR 2020poster

OAR segmentation is a critical step in radiotherapy of head and neck (H&N) cancer, where inconsistencies across radiation oncologists and prohibitive labor costs motivate automated approaches. However, leading methods using standard fully convolutional network workflows that are challenged when the…

Cited by 87PDFScholar
2020

Structured Landmark Detection via Topology-Adapting Deep Graph Learning

ECCV 2020poster

Image landmark detection aims to automatically identify the locations of predefined fiducial points. Despite recent success in this field, higher-ordered structural modeling to capture implicit or explicit relationships among anatomical landmarks has not been adequately exploited. In this work, we p…

Cited by 121SourcePDFScholar
2019

Adversarial Discrete Sequence Generation without Explicit NeuralNetworks as Discriminators

AISTATS 2019poster

This paper presents a novel approach to train GANs for discrete sequence generation without resorting to an explicit neural network as the discriminator. We show that when an alternative mini-max optimization procedure is performed for the value function where a closed form solution for the discrimi…

2019

An Autonomous Loop-Closure Approach for Simultaneous Exploration and Coverage of Unknown Infrastructure Using MAVs

ICRA 2019poster

The recent proliferation of low-cost Micro Aerial Vehicles (MAV) offers an attractive means for inspecting critical infrastructure autonomously. However, to enable such autonomous tasks requires a precise spatial model of the structure and operational area, typically constructed using sensor measure…

Cited by 16SourceScholar
2019

Long Term Background Reference Based Satellite Video Coding

ICASSP 2019accepted

Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellite calls for higher coding efficiency. In…

Cited by 0SourceScholar
2019

Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge

ICASSP 2019accepted

The rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy…

Cited by 0SourceScholar
2018

Cross-Modal Learning to Rank with Adaptive Listwise Constraint

ICASSP 2018accepted

Multi-modal data lies on heterogeneous feature spaces, which brings a significant challenge to cross-modal retrieval. Some works have been proposed to cope with this problem by learning a common subspace. However, previous methods often learn the common subspace by enhancing the relation between emb…

Cited by 0SourceScholar
2017

Shape-based object classification and recognition through continuum manipulation

IROS 2017poster

We introduce a novel approach to shape-based object classification and recognition through the use of a continuum manipulator. Noticing the fact that when a continuum manipulator wraps around an object in a whole-arm grasping, its own shape is indicative of the shape of the object, our approach enab…

Cited by 10SourceScholar