Auxiliary Pooling Layer For Spoken Language Understanding
Yukun Ma, Trung Hieu Nguyen, Jinjie Ni, Wen Wang, Qian Chen, Chong Zhang, Bin Ma
Abstract
End-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to encode semantic information. Existing works explore transferring knowledge from a pre-trained text model with different alignment losses at a fixed granularity. In this paper, we address the variable granularity in transferring knowledge from texts to speech representation via APLY, an auxiliary pooling layer, that fuses the global information with the adaptively encoded local context. We demonstrate the effectiveness of APLY on three benchmarks of spoken language understanding.
BibTeX
@inproceedings{icassp2023_auxiliarypooling,
title = {Auxiliary Pooling Layer For Spoken Language Understanding},
author = {Yukun Ma and Trung Hieu Nguyen and Jinjie Ni and Wen Wang and Qian Chen and Chong Zhang and Bin Ma},
booktitle = {ICASSP 2023},
year = {2023}
}