← Search

Xiaomei Zhang

11 accepted papers

2026

AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned Adaptation

AAAI 2026technical

The surface pressure field of transportation systems, including cars, trains, and aircraft, is critical for aerodynamic analysis and design. In recent years, deep neural networks have emerged as promising and efficient methods for modeling surface pressure field, being alternatives to computationall

Cited by 0SourcePDFScholar
2026

Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

ICLR 2026poster

Generating 3D human poses from multimodal inputs such as text or images requires models to capture both rich semantic and spatial correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise, their supervised fine-tuning (SFT) paradigm struggles to resolve the tas…

Cited by 0SourceScholar
2025

Diffusion Models are Zero-Shot Generative Text-Vision Retrievers

ICASSP 2025accepted

Large-scale text-to-image diffusion models have demonstrated impressive capabilities for downstream tasks by leveraging strong vision-language alignment from generative pre-training. Recently, a number of works have explored how to use the power of text-to-image diffusion models for text-image match…

Cited by 0SourceScholar
2025

MVBoost: Boost 3D Reconstruction with Multi-View Refinement

CVPR 2025poster

Recent advancements in 3D object reconstruction have been remarkable, yet most current 3D models rely heavily on existing 3D datasets. The scarcity of diverse 3D datasets results in limited generalization capabilities of 3D reconstruction models. In this paper, we propose a novel framework for boost…

2025

Macaque-Motion-Monitor Dataset: A New Benchmark for Macaque Action Recognition

ICASSP 2025accepted

Recent advancements in computational techniques significantly impact bioengineering, particularly in drug safety assessments and neuroscience trials using primate models. Macaques are extensively used due to their genetic and physiological similarities to humans. However, there is a scarcity of maca…

Cited by 0SourceScholar
2024

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

CVPR 2024poster

Achieving high synchronization in the synthesis of realistic speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity while Neural Radiance Fields (NeRF) methods although they can address thi…

2023

Graphics Capsule: Learning Hierarchical 3D Face Representations From 2D Images

CVPR 2023poster

The function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural ne…

Cited by 7SourcePDFScholar
2023

High-Fidelity Clothed Avatar Reconstruction From a Single Image

CVPR 2023poster

This paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR)…

2022

HP-Capsule: Unsupervised Face Part Discovery by Hierarchical Parsing Capsule Network

CVPR 2022poster

Capsule networks are designed to present the objects by a set of parts and their relationships, which provide an insight into the procedure of visual perception. Although recent works have shown the success of capsule networks on simple objects like digits, the human faces with homologous structures…

Cited by 22PDFScholar