← Search

Jiahang Cai

2 accepted papers

2026

Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation

CVPR 2026

Vision Transformers (ViTs) have recently achieved state-of-the-art performance in 2D human pose estimation due to their strong global modeling capability. However, existing ViT-based pose estimators are designed for static images and process each frame independently, thereby ignoring the temporal co

Cited by 0SourcecodeScholar
2026

End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer

AAAI 2026technical

Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single-person pose estimation. This design relies on heuristic operations such as tracking, RoI cropping, and non-maximum suppression, limi

Cited by 0SourcePDFScholar