← Search

Fan Ma

28 accepted papers

2026

AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows

CVPR 2026

Training-free 3D editing aims to modify 3D shapes based on human instructions without model finetuning. It plays a crucial role in 3D content creation. However, existing approaches often struggle to produce strong or geometrically stable edits, largely due to inconsistent latent anchors introduced b

Cited by 0SourcecodeScholar
2026

ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation

ICLR 2026poster

Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserving the identity of multiple distinct subjects. To address these limitations, we introduce **ContextGen**, a novel Diffu…

Cited by 0SourcecodeScholar
2026

Echoes of Ownership: Adversarial-Guided Dual Injection for Copyright Protection in MLLMs

CVPR 2026

With the rapid deployment of multimodal large language models (MLLMs), disputes regarding model ownership have become increasingly frequent, raising significant concerns about intellectual property protection. In this paper, we propose a framework for generating copyright triggers for MLLMs, enablin

Cited by 0SourcecodeScholar
2026

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

CVPR 2026

Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal control in multi-event scenarios or when visual cues are insufficient, such as small regions, off-screen sounds, or occlud

Cited by 0SourcecodeScholar
2026

LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization

ICLR 2026poster

Generating coherent and communicative visual sequences, such as image sequences and videos, remains a significant challenge for current multimodal systems. Despite advances in visual quality and the integration of world knowledge, existing models still struggle to maintain logical flow, often result…

Cited by 0SourceScholar
2025

Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion

AAAI 2025technical

Human motion generative models have enabled promising applications, but the ability of text-to-motion (T2M) models to produce realistic motions raises security concerns if exploited maliciously. Despite growing interest in T2M, limited research focus on safeguarding these models against adversarial…

Cited by 2SourcePDFScholar
2025

BrainGuard: Privacy-Preserving Multisubject Image Reconstructions from Brain Activities

AAAI 2025technical

Reconstructing perceived images from human brain activity forms a crucial link between human and machine learning through Brain-Computer Interfaces. Early methods primarily focused on training separate models for each individual to account for individual variability in brain activity, overlooking va…

2025

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

ICML 2025poster

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human preferences, limiting their applicability and flexibility. To address these limit…

Cited by 7SourcePDFScholar
2025

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

ICCV 2025poster

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference, potentially omitting crucial visual information. To address the challe…

Cited by 0SourcePDFScholar
2025

Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models

AAAI 2025technical

Diffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become increasingly important. Current metrics focus on directly m…

2025

InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation

ICCV 2025poster

We present InfiniDreamer, a novel framework for generating human motions of arbitrary length. Existing methods typically produce only short sequences, limited by the scarcity of long-range motion data. To address this, InfiniDreamer first generates short sub-motions for each textual description, the…

Cited by 0SourcePDFScholar
2025

Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection

ICLR 2025poster

Visual instructions for long-horizon tasks are crucial as they intuitively clarify complex concepts and enhance retention across extended steps. Directly generating a series of images using text-to-image models without considering the context of previous steps results in inconsistent images, increa…

Cited by 0SourcePDFScholar
2025

Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion

CVPR 2025poster

Animatable head avatar generation typically requires extensive data for training. To reduce the data requirements, a natural solution is to leverage existing data-free static avatar generation methods, such as pre-trained diffusion models with score distillation sampling (SDS), which align avatars w…

2024

CapHuman: Capture Your Moments in Parallel Universes

CVPR 2024poster

We concentrate on a novel human-centric image synthesis task that is given only one reference facial photograph it is expected to generate specific individual images with diverse head positions poses facial expressions and illuminations in different contexts. To accomplish this goal we argue that ou…

2024

HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting

ECCV 2024poster

"Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars efficiently. Previous animatable head models like FLAME have…

2024

Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval

CVPR 2024poster

We study the zero-shot Composed Image Retrieval (ZS-CIR) task which is to retrieve the target image given a reference image and a description without training on the triplet datasets. Previous works generate pseudo-word tokens by projecting the reference image features to the text embedding space. H…

2024

LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse Kernels

CVPR 2024poster

Autonomous systems need to process large-scale sparse and irregular point clouds with limited compute resources. Consequently it is essential to develop LiDAR perception methods that are both efficient and effective. Although naively enlarging 3D kernel size can enhance performance it will also lead…

2024

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

CVPR 2024highlight

We present a Multi-Instance Generation (MIG) task simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions the task is to ensure that generated instances are accurately at the designated locations and…

2024

Psychometry: An Omnifit Model for Image Reconstruction from Human Brain Activity

CVPR 2024poster

Reconstructing the viewed images from human brain activity bridges human and computer vision through the Brain-Computer Interface. The inherent variability in brain function between individuals leads existing literature to focus on acquiring separate models for each individual using their respective…

Cited by 17SourcePDFScholar
2024

Stitching Segments and Sentences towards Generalization in Video-Text Pre-training

AAAI 2024technical

Video-language pre-training models have recently achieved remarkable results on various multi-modal downstream tasks. However, most of these models rely on contrastive learning or masking modeling to align global features across modalities, neglecting the local associations between video frames and…

Cited by 6SourcePDFScholar
2024

VISTA-LLAMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

CVPR 2024poster

Recent advances in large video-language models have displayed promising outcomes in video comprehension. Current approaches straightforwardly convert video into language tokens and employ large language models for multi-modal tasks. However this method often leads to the generation of irrelevant con…

Cited by 16SourcePDFScholar
2024

VividDreamer: Invariant Score Distillation for Hyper-Realistic Text-to-3D Generation

ECCV 2024poster

"This paper presents Invariant Score Distillation (ISD), a novel method for high-fidelity text-to-3D generation. ISD aims to tackle the over-saturation and over-smoothing problems in Score Distillation Sampling (SDS). In this paper, SDS is decoupled into a weighted sum of two components: the reconst…

2022

Unified Transformer Tracker for Object Tracking

CVPR 2022poster

As an important area in computer vision, object tracking has formed two separate communities that respectively study Single Object Tracking (SOT) and Multiple Object Tracking (MOT). However, current methods in one tracking scenario are not easily adapted to the other due to the divergent training da…

Cited by 138PDFcodeScholar
2020

SF-Net: Single-Frame Supervision for Temporal Action Localization

ECCV 2020poster

In this paper, we study an intermediate form of supervision, i.e., single-frame supervision, for temporal action localization (TAL). To obtain the single-frame supervision, the annotators are asked to identify only a single frame within the temporal window of an action. This can significantly reduce…