← Search

Jixuan Wang

8 accepted papers

2023

End-to-End Spoken Language Understanding Using Joint CTC Loss and Self-Supervised, Pretrained Acoustic Encoders

ICASSP 2023accepted

It is challenging to extract semantic meanings directly from audio signals in spoken language understanding (SLU), due to the lack of textual information. Popular end-to-end (E2E) SLU models utilize sequence-to-sequence automatic speech recognition (ASR) models to extract textual embeddings as input…

Cited by 0SourceScholar
2023

Pyramid Dynamic Inference: Encouraging Faster Inference Via Early Exit Boosting

ICASSP 2023accepted

Transformer-based models demonstrate state of the art results on several natural language understanding tasks. However, their deployment comes at the cost of increased footprint and inference latency, limiting their adoption to real-time applications. Early exit strategies are designed to speed-up t…

Cited by 0SourceScholar
2023

Quantifying Catastrophic Forgetting in Continual Federated Learning

ICASSP 2023accepted

The deployment of Federated Learning (FL) systems poses various challenges such as data heterogeneity and communication efficiency. We focus on a practical FL setup that has recently drawn attention, where the data distribution on each device is not static but dynamically evolves over time. This set…

Cited by 0SourceScholar
2021

Encoding Syntactic Knowledge in Transformer Encoder for Intent Detection and Slot Filling

AAAI 2021technical

We propose a novel Transformer encoder-based architecture with syntactical knowledge encoded for intent detection and slot filling. Specifically, we encode syntactic knowledge into the Transformer encoder by jointly training it to predict syntactic parse ancestors and part-of-speech of each token vi…

Cited by 42SourcePDFScholar
2021

Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation

NeurIPS 2021poster

Large pretrained language models (LMs) like BERT have improved performance in many disparate natural language processing (NLP) tasks. However, fine tuning such models requires a large number of training examples for each target task. Simultaneously, many realistic NLP problems are "few shot", withou…

2020

Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural Networks

ICASSP 2020accepted

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be suboptimal for distinguishing speakers locally in a specific meet…

Cited by 0SourceScholar
2019

Centroid-based Deep Metric Learning for Speaker Recognition

ICASSP 2019accepted

Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set a…

Cited by 0SourceScholar