2024
Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion
ECCV 2024poster
"This paper presents a self-supervised temporal video alignment framework which is useful for several fine-grained human activity understanding applications. In contrast with the state-of-the-art method of CASA, where sequences of 3D skeleton coordinates are taken directly as input, our key idea is…