Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical features
Taufiq Hasan, Mohammed Abdel-Wahab, Srinivas Parthasarathy, Carlos Busso, Yang Liu
Abstract
Research on automatic speech summarization typically focuses on optimizing objective evaluation criteria, such as the ROUGE metric, which depend on word and phrase overlaps between automatic and manually generated summary documents. However, the actual quality of the speech summarizer largely depends on how the end-users perceive the audio output. This work focuses on the task of composing summarized audio streams with the aim of improving the quality and interest perceived by the end-user. First, using crowd-sourced summary annotations on a broadcast news corpus, we train a rank-SVM classifier to learn the relative importance of each sentence in a news story. Acoustic, lexical and structural features are used for training. In addition, we investigate the perceived emotion level in each sentence to aid the summarizer in selecting interesting sentences, yielding an emotion-aware summarizer. Next, we propose several methods to combine these sentences to generate a compressed audio stream. Subjective evaluations are performed to evaluate the quality of the generated summaries on the following criterion: interest, abruptness, informativeness, attractiveness, and overall quality. The results indicate that users are most sensitive to the linguistic coherence and continuity of the audio stream.
BibTeX
@inproceedings{icassp2016_automaticcomposi,
title = {Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical features},
author = {Taufiq Hasan and Mohammed Abdel-Wahab and Srinivas Parthasarathy and Carlos Busso and Yang Liu},
booktitle = {ICASSP 2016},
year = {2016}
}