ACL 2022long105 citations

SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities

Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang, Kushal Lakhotia, Shu-wen Yang, Shuyan Dong, Andy Liu

Abstract

Transfer learning has proven to be crucial in advancing the state of speech and natural language processing research in recent years. In speech, a model pre-trained by self-supervised learning transfers remarkably well on multiple tasks. However, the lack of a consistent evaluation methodology is limiting towards a holistic understanding of the efficacy of such models. SUPERB was a step towards introducing a common benchmark to evaluate pre-trained models across various speech tasks. In this paper, we introduce SUPERB-SG, a new benchmark focusing on evaluating the semantic and generative capabilities of pre-trained models by increasing task diversity and difficulty over SUPERB. We use a lightweight methodology to test the robustness of representations learned by pre-trained models under shifts in data domain and quality across different types of tasks. It entails freezing pre-trained model parameters, only using simple task-specific trainable heads. The goal is to be inclusive of all researchers, and encourage efficient use of computational resources. We also show that the task diversity of SUPERB-SG coupled with limited task supervision is an effective recipe for evaluating the generalizability of model representation.

BibTeX
@inproceedings{tsai-etal-2022-superb,
    title = "{SUPERB}-{SG}: Enhanced Speech processing Universal {PER}formance Benchmark for Semantic and Generative Capabilities",
    author = "Tsai, Hsiang-Sheng  and
      Chang, Heng-Jui  and
      Huang, Wen-Chin  and
      Huang, Zili  and
      Lakhotia, Kushal  and
      Yang, Shu-wen  and
      Dong, Shuyan  and
      Liu, Andy  and
      Lai, Cheng-I  and
      Shi, Jiatong  and
      Chang, Xuankai  and
      Hall, Phil  and
      Chen, Hsuan-Jui  and
      Li, Shang-Wen  and
      Watanabe, Shinji  and
      Mohamed, Abdelrahman  and
      Lee, Hung-yi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.580/",
    doi = "10.18653/v1/2022.acl-long.580",
    pages = "8479--8492"
}
SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities · ACL 2022