← Search

Wanze Xie

2 accepted papers

2022

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

NeurIPS 2022accept

Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional,…

Cited by 23SourcePDFScholar
2021

MOMA: Multi-Object Multi-Actor Activity Parsing

NeurIPS 2021poster

Complex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Recognizing activities poses a challenge that requires a detailed understanding of a…

Cited by 32SourcePDFScholar