Denoising 3D Human Poses from Low-Resolution Video using Variational Autoencoder
Chihiro Nakatsuka, Satoshi Komorita
Abstract
We tackle the problem of refining and denoising a series of 3D human poses estimated from a low-resolution video. Low-resolution often causes the wrong pose estimation, e.g., left-right switching and the absence of keypoints. We propose to use the variational autoencoder (VAE) to remove these challenging noises. The VAE model utilizes time-series information and motion priors in denoising. From our experiments, the VAE model can reduce the pose estimation error (MPJPE) for poor-quality images by 24.37mm, from the original 105.53mm. This improves about 6.5 times over the traditional DCT approach. In addition, it removes jitters and generates smooth movements, which is helpful in recognition of human behaviors.
BibTeX
@inproceedings{iros2021_denoising3dhuman,
title = {Denoising 3D Human Poses from Low-Resolution Video using Variational Autoencoder},
author = {Chihiro Nakatsuka and Satoshi Komorita},
booktitle = {IROS 2021},
year = {2021}
}