← Search

Chongyang Wang

3 accepted papers

2025

Multi-View 3D Human Pose Estimation with Weakly Synchronized Images

AAAI 2025technical

Multi-view 3D human pose estimation (MHPE) is an important research task in computer vision. To maintain consistency during the data collection, hardware synchronization devices are commonly used to connect cameras, ensuring that images from different views are captured simultaneously. However, sync…

Cited by 0SourcePDFScholar
2023

Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer

AAAI 2023technical

This paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors. We achieve this with a pyramidal multimodal transformer (PMT) model, which simply incorporates a learnable word embedd…

2022

Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering

IJCAI 2022poster

Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore the visual appearance-motion information at different temporal scales, it is unknown how to incorporate the multilevel pro…