Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing
Ross Cutler, Ramin Mehran, Sam Johnson, Cha Zhang, Adam Kirk, Oliver Whyte, Adarsh Kowdle
Abstract
Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a new automated ASD and VC that performs within 0.3 MOS of an expert cinematographer based on subjective ratings with a 1-5 scale. This system uses a 4K wide-FOV camera, a depth camera, and a microphone array, extracts features from each modality and trains an ASD using an AdaBoost machine learning system that is very efficient and runs in real-time. A VC is similarly trained using machine learning. To avoid distracting the room participants the system has no moving parts - the VC works by cropping and zooming the 4K wide-FOV video stream. The system was tuned and evaluated using extensive crowdsourcing techniques and evaluated on a system with N=100 meetings, each 25 minutes in length.
BibTeX
@inproceedings{icassp2020_multimodalactive,
title = {Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing},
author = {Ross Cutler and Ramin Mehran and Sam Johnson and Cha Zhang and Adam Kirk and Oliver Whyte and Adarsh Kowdle},
booktitle = {ICASSP 2020},
year = {2020}
}