← Search

Yue Ma

39 accepted papers

2026

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization

CVPR 2026

We introduce AvatarPointillist, a novel framework for generating dynamic 4D Gaussian avatars from a single portrait image. At the core of our method is a decoder-only Transformer that autoregressively generates a point cloud for 3D Gaussian Splatting. This sequential approach allows for precise, ada

Cited by 0SourcecodeScholar
2026

CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing

CVPR 2026

Unified diffusion editors often rely on a fixed, shared backbone for diverse tasks, suffering from task interference and poor adaptation to heterogeneous demands (e.g., local vs global, semantic vs photometric). In particular, prevalent ControlNet and OmniControl variants combine multiple conditioni

Cited by 0SourcecodeScholar
2026

EasyCreator: Empowering 4D Creation through Video Inpainting

ICLR 2026poster

We introduce EasyCreator, a novel 4D video creation framework capable of both generating and editing 4D content from a single monocular video input. By leveraging a powerful video inpainting foundation model as a generative prior, we reformulate 4D video creation as a video inpainting task, enabling…

Cited by 0SourceScholar
2026

EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

ICLR 2026poster

Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoRAs) finetuning to obtain better performance. However, existing adaptation-based…

Cited by 0SourceScholar
2026

FastVMT: Eliminating Redundancy in Video Motion Transfer

ICLR 2026poster

Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent methods predominantly use the Diffusion Transformer (DiT) architecture. To achieve satisfactory runtime, several methods…

Cited by 0SourceScholar
2026

Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control

ICLR 2026poster

While recent flow-based image editing models demonstrate general-purpose capabilities across diverse tasks, they often struggle to specialize in challenging scenarios---particularly those involving large-scale shape transformations. When performing such structural edits, these methods either fail t…

Cited by 0SourcecodeScholar
2026

Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers

AAAI 2026technical

Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature caching techniques have been proposed to accelerate inference by reusing hidden representations from previous timesteps. Howev

Cited by 0SourcePDFScholar
2026

Group Editing: Edit Multiple Images in One Go

CVPR 2026

In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images may vary significantly in pose, viewpoint, and spatial layout. Achieving coherent edits requires establishing reliable co

Cited by 9SourcecodeScholar
2026

HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration

ICLR 2026poster

Diffusion models have achieved remarkable success in content generation but suffer from prohibitive computational costs due to iterative sampling. While recent feature caching methods tend to accelerate inference through temporal extrapolation, these methods still suffer from severe quality loss due…

Cited by 0SourcecodeScholar
2026

Tea-Adapter: Teacher Adapter for Efficient Conditional Generation

CVPR 2026

We propose Tea-Adapter, a plug-and-play adapter designed to efficiently integrate conditional knowledge from a smaller teacher model into a larger student video diffusion model. Existing controllable video DiT methods face critical challenges: full fine-tuning of billion-parameter models is extremel

Cited by 0SourceScholar
2025

DiT4Edit: Diffusion Transformer for Image Editing

AAAI 2025technical

Despite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to effectively capture the long-range dependencies among patches, leading to highe…

2025

DuPI: Dual-resolution Pseudo-label Integration for Semi-supervised Instance Segmentation

ICASSP 2025accepted

The role of high-quality pseudo-labels is pivotal in semi-supervised instance segmentation (SSIS). However, existing SSIS frameworks predominantly produce pseudo-labels at a single resolution, which can introduce noise that adversely affects the quality of learning at both the pixel level and in ter…

Cited by 0SourceScholar
2025

EEdit : Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing

ICCV 2025poster

Inversion-based image editing is rapidly gaining momentum while suffering from significant computation overhead, hindering its application in real-time interactive scenarios. In this paper, we rethink that the redundancy in inversion-based image editing exists in both the spatial and temporal dimens…

2025

Enhancing Image Generation Fidelity via Progressive Prompts

ICASSP 2025accepted

Diffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propos…

Cited by 0SourceScholar
2025

Follow-Your-Click: Open-domain Regional Image Animation via Motion Prompts

AAAI 2025technical

Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. A…

Cited by 52SourcePDFScholar
2025

Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance

ICASSP 2025accepted

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently…

Cited by 0SourceScholar
2025

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving

AAAI 2025technical

With the prevalence of Multimodal Large Language Models(MLLMs), autonomous driving has encountered new opportunities and challenges. In particular, multi-modal video understanding is critical to interactively analyze what will happen in the procedure of autonomous driving. However, videos in such a…

Cited by 2SourcePDFScholar
2025

IMU-Based Motion Mode Recognition in Soft Underwater Exosuit

IROS 2025

By accurately recognizing the wearer’s motion, the underwater exoskeleton enables more efficient human-machine collaboration and provides enhanced assistance in complex and dynamic underwater environments. In this study, we propose a soft underwater exosuit motion mode recognizer based on a long sho

Cited by 0SourceScholar
2025

Infinite-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation

AAAI 2025technical

This paper explores higher-resolution video outpainting with extensive content generation. We point out common issues faced by existing methods when attempting to largely outpaint videos: the generation of low-quality content and limitations imposed by GPU memory. To address these challenges, we pro…

2025

InstantSwap: Fast Customized Concept Swapping across Sharp Shape Differences

ICLR 2025poster

Recent advances in Customized Concept Swapping (CCS) enable a text-to-image model to swap a concept in the source image with a customized target concept. However, the existing methods still face the challenges of $\textit{\textbf{inconsistency}}$ and $\textit{\textbf{inefficiency}}$. They struggle t…

2025

MultiBooth: Towards Generating All Your Concepts in an Image from Text

AAAI 2025technical

This paper introduces MultiBooth, a method that generates images from texts containing various concepts from users.Despite diffusion models bringing significant advancements for customized text-to-image generation, existing methods often struggle with multi-concept scenarios due to low concept fidel…

2025

RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging

EMNLP 2025

We unveil that internal representations in large language models (LLMs) serve as reliable proxies of learned knowledge, and propose **RECALL**, a novel representation-aware model merging framework for continual learning without access to historical data. RECALL computes inter-model similarity from l

2025

Taming Rectified Flow for Inversion and Editing

ICML 2025poster

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often struggle with inversion inaccuracies, which could further limit their effectivenes…

2025

Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

ICLR 2025poster

Controllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. T…

Cited by 0SourcePDFScholar
2025

Uncertainty-aware Probabilistic 3D Human Motion Forecasting via Invertible Networks

ICRA 2025

3D human motion forecasting aims to enable autonomous applications. Estimating uncertainty for each prediction (i.e., confidence based on probability density or quantile) is essential for safety-critical contexts like human-robot collaboration to minimize risks. However, existing diverse motion fore

Cited by 3SourceScholar
2025

Underwater Exosuit Actuator Design for Unrestricted Bidirectional Hip Assistance During Flutter Kicking

IROS 2025

Underwater assistance is crucial for individuals who depend on diving for their livelihood. In this paper, we propose a novel underwater exosuit actuator designed to assist with flutter kicking during diving, thereby decreasing the effort the diver has to exert. The actuator can provide bidirectiona

Cited by 0SourceScholar
2024

A Novel Cross-Sensor Self-Supervised Learning Method for Rotating Machinery Fault Diagnosis

ICASSP 2024accepted

Fault diagnosis is crucial in mechanical prognostics and health management. However, fault features extracted from single-sensor data are limited in complex operating environments. Extracting complementary and robust fault features from multi-sensor monitoring data is essential, especially under lim…

Cited by 0SourceScholar
2024

Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

CVPR 2024poster

Video Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as similar video grounding problems and address them together with transformer-based architecture. However we observe that the…

2024

COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing

NeurIPS 2024poster

Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite extensive efforts, maintaining the temporal consistency of edited videos remains challenging due to the lack of temporal…

2024

Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free Videos

AAAI 2024technical

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose captions and the generative prior models for videos. In this w…

2023

SemanticAC: Semantics-Assisted Framework for Audio Classification

ICASSP 2023accepted

In this paper, we propose SemanticAC, a semantics-assisted framework for Audio Classification to better leverage the semantic information. Unlike conventional audio classification methods that treat class labels as discrete vectors, we employ a language model to extract abundant semantics from label…

Cited by 0SourceScholar