ICASSP 2018accepted0 citations

Improved Audio-Visual Laughter Detection Via Multi-Scale Multi-Resolution Image Texture Features and Classifier Fusion

Zahid Akhtar, Stefany Bedoya, Tiago H. Falk

Abstract

Efforts are afoot to design better context-aware human-computer interaction techniques that have knowledge of both their surrounding and the affective state of the user. One of the most important nonverbal behavioural cues for affective human-machine interaction is laughter. Automatic detection of laughter is an interesting, yet challenging problem, which in recent years has gained increased attention from both the academic and industrial communities. The majority of existing laughter detection systems rely on either audio or video modalities. Humans, however, typically rely on audio-visual cues during conversation and/or interaction, thus it is expected that improved results can be achieved if both modalities are used. In this work, we propose a multimodal framework that analyzes audio and video channels separately, then fuses their decisions. Conventional speech spectral and prosodic features are used, whereas new multi -scale multiresolution binarized statistical image features are proposed due to their improved expressive power. Experiments with the publicly available MAHNOB Laughter database show that decision level fusion based on support vector machine classifiers leads to improved performance over single modality approaches, as well as over previously-proposed methods, all whilst requiring just a fraction of the computational power.

BibTeX
@inproceedings{icassp2018_improvedaudiovis,
  title = {Improved Audio-Visual Laughter Detection Via Multi-Scale Multi-Resolution Image Texture Features and Classifier Fusion},
  author = {Zahid Akhtar and Stefany Bedoya and Tiago H. Falk},
  booktitle = {ICASSP 2018},
  year = {2018}
}