ICASSP 2023accepted0 citations

I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech Recognition

Yifan Peng, Jaesong Lee, Shinji Watanabe

Abstract

Transformer-based end-to-end speech recognition has achieved great success. However, the large footprint and computational overhead make it difficult to deploy these models in some real-world applications. Model compression techniques can reduce the model size and speed up inference, but the compressed model has a fixed architecture which might be suboptimal. We propose a novel Transformer encoder with Input-Dependent Dynamic Depth (I3D) to achieve strong performance-efficiency trade-offs. With a similar number of layers at inference time, I3D-based models outperform the vanilla Transformer and the static pruned model via iterative layer pruning. We also present interesting analysis on the gate probabilities and the input-dependency, which helps us better understand deep encoders.

BibTeX
@inproceedings{icassp2023_i3dtransformerar,
  title = {I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech Recognition},
  author = {Yifan Peng and Jaesong Lee and Shinji Watanabe},
  booktitle = {ICASSP 2023},
  year = {2023}
}
I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech Recognition · ICASSP 2023