"PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation"
"Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of human-centric vision, albeit CLIP-like representations encode…