AAAI 2024technical5 citations

Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting

Qi Zhang, Yunfei Gong, Daijie Chen, Antoni B. Chan, Hui Huang

Abstract

Recent deep learning-based multi-view people detection (MVD) methods have shown promising results on existing datasets. However, current methods are mainly trained and evaluated on small, single scenes with a limited number of multi-view frames and fixed camera views. As a result, these methods may not be practical for detecting people in larger, more complex scenes with severe occlusions and camera calibration errors. This paper focuses on improving multi-view people detection by developing a supervised view-wise contribution weighting approach that better fuses multi-camera information under large scenes. Besides, a large synthetic dataset is adopted to enhance the model's generalization ability and enable more practical evaluation and comparison. The model's performance on new testing scenes is further improved with a simple domain adaptation technique. Experimental results demonstrate the effectiveness of our approach in achieving promising cross-scene multi-view people detection performance.

BibTeX
@article{Zhang_Gong_Chen_Chan_Huang_2024, title={Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/28553}, DOI={10.1609/aaai.v38i7.28553}, abstractNote={Recent deep learning-based multi-view people detection (MVD) methods have shown promising results on existing datasets. However, current methods are mainly trained and evaluated on small, single scenes with a limited number of multi-view frames and fixed camera views. As a result, these methods may not be practical for detecting people in larger, more complex scenes with severe occlusions and camera calibration errors. This paper focuses on improving multi-view people detection by developing a supervised view-wise contribution weighting approach that better fuses multi-camera information under large scenes. Besides, a large synthetic dataset is adopted to enhance the model’s generalization ability and enable more practical evaluation and comparison. The model’s performance on new testing scenes is further improved with a simple domain adaptation technique. Experimental results demonstrate the effectiveness of our approach in achieving promising cross-scene multi-view people detection performance.}, number={7}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Zhang, Qi and Gong, Yunfei and Chen, Daijie and Chan, Antoni B. and Huang, Hui}, year={2024}, month={Mar.}, pages={7242-7250} }