← Search

Xiaoqiang Liu

9 accepted papers

2026

3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation

CVPR 2026

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel-view synthesis. Explicit 3D models, though structurally i

Cited by 0SourcecodeScholar
2026

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

ICML 2026poster

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spa…

Cited by 0SourceScholar
2025

Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control

ICLR 2025poster

Speech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting flexible fine-grained facial control within the spatiotemporal d…

Cited by 0SourcePDFScholar
2025

Dynamic Routing and Calibration for Few-Shot Object Detection

ICASSP 2025accepted

Few-shot object detection (FSOD), aiming to enhance the performance of novel object detection with limited labeled samples, has recently gained significant attention. Recent researches primarily focus on improving the generalization of novel classes and enhancing detector performance. However, the d…

Cited by 0SourceScholar
2025

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

ICCV 2025poster

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head rotations and out-of-distribution (OOD) audio. Moreover, they are…

Cited by 0SourcePDFScholar
2025

GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian Projections

CVPR 2025poster

Existing radiance field-based head avatar methods have mostly relied on pre-computed explicit priors (e.g., mesh, point) or neural implicit representations, making it challenging to achieve high fidelity with both computational efficiency and low memory consumption. To overcome this, we present GPAv…

Cited by 0SourcePDFScholar
2025

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

ICML 2025spotlight

Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, p…

Cited by 0SourcePDFScholar
2025

OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers

NeurIPS 2025spotlight

Lip synchronization is the task of aligning a speaker’s lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames and masked-frame inpainting, which limit their robustness to…

Cited by 0SourceScholar
2024

Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question Answering

AAAI 2024technical

Few-shot Visual Question Answering (VQA) realizes few-shot cross-modal learning, which is an emerging and challenging task in computer vision. Currently, most of the few-shot VQA methods are confined to simply extending few-shot classification methods to cross-modal tasks while ignoring the spatial…

Cited by 3SourcePDFScholar