← Search

Md. Akmal Haidar

8 accepted papers

2022

CILDA: Contrastive Data Augmentation Using Intermediate Layer Knowledge Distillation

COLING 2022main

Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by leveraging Contrastive Learning, Intermediate Layer Distillation, Data Augmentation, and Adversarial Training. In this wor…

Cited by 4SourcePDFScholar
2022

RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation

NAACL 2022findings

Intermediate layer knowledge distillation (KD) can improve the standard KD technique (which only targets the output of teacher and student models) especially over large pre-trained language models. However, intermediate layer distillation suffers from excessive computational burdens and engineering…

Cited by 26SourcePDFScholar
2021

Fine-Tuning of Pre-Trained End-to-End Speech Recognition with Generative Adversarial Networks

ICASSP 2021accepted

Adversarial training of end-to-end (E2E) ASR systems using generative adversarial networks (GAN) has recently been explored for low-resource ASR corpora. GANs help to learn the true data representation through a two-player min-max game. However, training an E2E ASR model using a large ASR corpus wit…

Cited by 0SourceScholar
2021

Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge Distillation

EMNLP 2021main

Intermediate layer matching is shown as an effective approach for improving knowledge distillation (KD). However, this technique applies matching in the hidden spaces of two different networks (i.e. student and teacher), which lacks clear interpretability. Moreover, intermediate layer KD cannot easi…

Cited by 29SourcePDFScholar
2020

From Unsupervised Machine Translation to Adversarial Text Generation

ICASSP 2020accepted

We present a self-attention based bilingual adversarial text generator (B-GAN) which can learn to generate text from the encoder representation of an unsupervised neural machine translation system. B-GAN is able to generate a distributed latent space representation which can be paired with an attent…

Cited by 0SourceScholar
2018

Reg-Gan: Semi-Supervised Learning Based on Generative Adversarial Networks for Regression

ICASSP 2018accepted

This research concerns introducing a method to solve the semi-supervised learning problem with generative adversarial networks (GANs) for regression. In contrast to classification, where only a limited number of distinct classes is given, the regression task is defined as predicting continuous label…

Cited by 0SourceScholar
2017

LDA-based context dependent recurrent neural network language model using document-based topic distribution of words

ICASSP 2017accepted

Adding context information into recurrent neural network language models (RNNLMs) have been investigated recently to improve the effectiveness of learning RNNLM. Conventionally, a fast approximate topic representation for a block of words was proposed by using corpus-based topic distribution of word…

Cited by 0SourceScholar