← Search

Mingtao Pei

10 accepted papers

2026

Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining

CVPR 2026

Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models' reliance on 3D training data from constrained environments li

Cited by 0SourceScholar
2025

PrimHOI: Compositional Human-Object Interaction via Reusable Primitives

ICCV 2025accepted

Synthesizing realistic Human-Object Interaction (HOI) motions is essential for creating believable digital characters and intelligent robots. Existing approaches rely on data-intensive learning models that struggle with the compositional structure of daily HOI motions, particularly for complex multi…

Cited by 0SourcePDFScholar
2025

Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing

ICCV 2025poster

We address the challenge of lifting 2D visual segmentation to 3D in Gaussian Splatting. Existing methods often suffer from inconsistent 2D masks across viewpoints and produce noisy segmentation boundaries as they neglect these semantic cues to refine the learned Gaussians. To overcome this, we intro…

Cited by 0SourcePDFScholar
2024

Automatic Radiology Reports Generation via Memory Alignment Network

AAAI 2024technical

The automatic generation of radiology reports is of great significance, which can reduce the workload of doctors and improve the accuracy and reliability of medical diagnosis and treatment, and has attracted wide attention in recent years. Cross-modal mapping between images and text, a key component…

Cited by 11SourcePDFScholar
2023

Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering

CVPR 2023poster

Video Question Answering (VideoQA) is challenging as it requires capturing accurate correlations between modalities from redundant information. Recent methods focus on the explicit challenges of the task, e.g. multimodal feature extraction, video-text alignment and fusion. Their frameworks reason th…

2022

Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports Generation

AAAI 2022technical

In this paper, we propose a vision-language pre-training model, Clinical-BERT, for the medical domain, and devise three domain-specific tasks: Clinical Diagnosis (CD), Masked MeSH Modeling (MMM), Image-MeSH Matching (IMM), together with one general pre-training task: Masked Language Modeling (MLM),…