← Search

Chenguang Ma

10 accepted papers

2026

EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human Animation

AAAI 2026technical

Recent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computational demands. Moreover, traditional work typically employs separate models for

Cited by 0SourcePDFScholar
2026

Exposing and Evaluating Hallucinations for GUI Grounding

CVPR 2026

Existing GUI benchmarks primarily focus on evaluating models' comprehensive capabilities but largely overlook hallucination phenomena in grounding tasks, which are crucial to the reliability of GUI understanding. In this work, we expose two major types of hallucinations in GUI grounding: 1) Confusio

Cited by 0SourceScholar
2026

M$^2$-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining

ICLR 2026poster

Graphical User Interface (GUI) agent is pivotal to advancing intelligent human-computer interaction paradigms. Constructing powerful GUI agents necessitates the large-scale annotation of high-quality user-behavior trajectory data (\textit{i.e.}, intent–trajectory pairs) for training. However, manual…

Cited by 0SourceScholar
2026

QuietPrune: Query-Guided Early Token Pruning for Vision-Language Models

CVPR 2026

Vision-language models (VLMs) demonstrate powerful capabilities in multimodal tasks. However, the large number of visual tokens imposes a significant computational cost. In this paper, we propose QuietPrune, a QUery-guIded Early Token Pruning method to remove redundant visual tokens within VLMs, the

Cited by 0SourcecodeScholar
2026

SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM Quantization

AAAI 2026technical

Deploying Vision-Language Models (VLMs) on edge devices (e.g., smartphones and robots) is crucial for enabling low-latency and privacy-preserving intelligent applications. Given the resource constraints of these devices, quantization offers a promising solution by improving memory efficiency and red

Cited by 0SourcePDFScholar
2025

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

AAAI 2025technical

The area of portrait image animation, propelled by audio input, has witnessed notable progress in the generation of lifelike and dynamic portraits. Conventional methods are limited to utilizing either audios or facial key points to drive images into videos, while they can yield satisfactory results,…

2025

EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation

CVPR 2025poster

Recent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality. However, these methods often face practical challenges due to extra control conditions, cumbersome condition injection modules, or limitation to head region driving. He…

2025

Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency

CVPR 2025poster

As a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-world online videos are often plagued by degradations such as blurring and quantization noise, due to the high compression ratio caused by high communication costs and limited tran…

2024

SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models

ECCV 2024poster

"Text-to-image diffusion models (SD) exhibit significant advancements while requiring extensive computational resources. Existing acceleration methods usually require extensive training and are not universally applicable. LCM-LoRA, trainable once for diverse models, offers universality but rarely co…

2022

FaceVerse: A Fine-Grained and Detail-Controllable 3D Face Morphable Model From a Hybrid Dataset

CVPR 2022poster

We present FaceVerse, a fine-grained 3D Neural Face Model, which is built from hybrid East Asian face datasets containing 60K fused RGB-D images and 2K high-fidelity 3D head scan models. A novel coarse-to-fine structure is proposed to take better advantage of our hybrid dataset. In the coarse module…

Cited by 110PDFcodeScholar