Dynamic ROI Adaptation for Accurate Non-Contact Heart Rate Estimation Using VGG-13 based Encoder-Decoder Model and Facial Landmarks
Aravind A Anil, Srinivasa Karthik, Mohanasankar Sivaprakasam, Jayaraj Joseph
Abstract
Heart rate monitoring using cameras is gaining attention for unobtrusive health tracking. This paper proposes an improved method for HR calculation using ROI adaptation tailored to individual facial features. A VGG-13 based encoder-decoder model was used for skin segmentation, classifying image pixels into left cheek, right cheek, and forehead regions, achieving a Dice coefficient of 0.90. ROIs were defined by the exterior coordinates of these skin pixels, adapting naturally to features like forehead size, cheek size, hairstyles, and spectacles. To address the challenge of ROIs with insufficient skin pixel coverage, we integrated Google MediaPipe landmarks to refine ROI selection, ensuring that at least 30 percent of the pixels in each region were skin pixels. The adaptive multi-ROI method was compared with traditional single-forehead and non-adaptive multi-ROI methods using the POS and CHROM rPPG algorithms. The adaptive approach showed better performance, achieving MAE values of 1.51 and 2.15 for POS and CHROM.
BibTeX
@inproceedings{icassp2025_dynamicroiadapta,
title = {Dynamic ROI Adaptation for Accurate Non-Contact Heart Rate Estimation Using VGG-13 based Encoder-Decoder Model and Facial Landmarks},
author = {Aravind A Anil and Srinivasa Karthik and Mohanasankar Sivaprakasam and Jayaraj Joseph},
booktitle = {ICASSP 2025},
year = {2025}
}