← Search

Xiang-Dong Zhou

4 accepted papers

2024

Deep Semantic Graph Transformer for Multi-View 3D Human Pose Estimation

AAAI 2024technical

Most Graph Convolutional Networks based 3D human pose estimation (HPE) methods were involved in single-view 3D HPE and utilized certain spatial graphs, existing key problems such as depth ambiguity, insufficient feature representation, or limited receptive fields. To address these issues, we propose…

2023

Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer

AAAI 2023technical

This paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors. We achieve this with a pyramidal multimodal transformer (PMT) model, which simply incorporates a learnable word embedd…

2022

Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering

IJCAI 2022poster

Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore the visual appearance-motion information at different temporal scales, it is unknown how to incorporate the multilevel pro…