← Search

Yuming Li

13 accepted papers

2026

AutoQVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization

ICLR 2026poster

The advent of Vision-Language-Action (VLA) models represents a significant leap for embodied intelligence, yet their immense computational demands critically hinder deployment on resource-constrained robotic platforms. Intuitively, low-bit quantization is a prevalent and preferred technique for larg…

Cited by 0SourcecodeScholar
2026

Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training

CVPR 2026

Recent advances in video generation models have achieved impressive results. However, these models heavily rely on the use of high-quality data that combines both high visual quality and high motion quality. In this paper, we identify a key challenge in video data curation: the Motion-Vision Quality

Cited by 0SourceScholar
2026

BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models

ICLR 2026poster

Recent progress in aligning image and video generative models with Group Relative Policy Optimization (GRPO) has improved human preference alignment, but existing variants remain inefficient due to sequential rollouts and large numbers of sampling steps, unreliable credit assignment,as sparse termin…

Cited by 0SourcecodeScholar
2026

EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human Animation

AAAI 2026technical

Recent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computational demands. Moreover, traditional work typically employs separate models for

Cited by 0SourcePDFScholar
2026

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

ICML 2026poster

Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14B parametered model typically demands hundreds of GPU days per experiment. Existing efficiency methods reduce costs thro…

Cited by 0SourceScholar
2026

ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory

AAAI 2026technical

Data scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availa

Cited by 0SourcePDFScholar
2026

ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance

ICASSP 2026poster

While recent advancements in robotic manipulation video synthesis have shown promise, significant challenges persist in ensuring effective instruction-following and achieving high visual quality. Recent methods, like RoboDreamer, utilize linguistic decomposition to divide instructions into separate…

Cited by 0SourcePDFScholar
2025

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

AAAI 2025technical

The area of portrait image animation, propelled by audio input, has witnessed notable progress in the generation of lifelike and dynamic portraits. Conventional methods are limited to utilizing either audios or facial key points to drive images into videos, while they can yield satisfactory results,…

2025

EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation

CVPR 2025poster

Recent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality. However, these methods often face practical challenges due to extra control conditions, cumbersome condition injection modules, or limitation to head region driving. He…

2025

Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency

CVPR 2025poster

As a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-world online videos are often plagued by degradations such as blurring and quantization noise, due to the high compression ratio caused by high communication costs and limited tran…

2025

Visual-Instructed Degradation Diffusion for All-in-One Image Restoration

CVPR 2025poster

Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that…

2016

Face liveness detection and recognition using shearlet based feature descriptors

ICASSP 2016accepted

Face recognition is a widely used biometric technology due to its convenience but it is vulnerable to spoofing attacks made by non-real faces such as a photograph or video of valid user. Face liveness detection is a core technology to make sure that the input face is a live person. However, this is…

Cited by 0SourceScholar
2015

Dynamic ROI based on K-means for remote photoplethysmography

ICASSP 2015accepted

Remote imaging photoplethysmography (RIPPG) can achieve contactless human vital signs monitoring. Though the remote operation mode brings a great convenience for RIPPG applications, the RIPPG signal quality is limited by the remote nature. Improving the RIPPG signal quality becomes an essential task…

Cited by 0SourceScholar