← Search

Ao Ma

22 accepted papers

2026

AutoPP: Towards Automated Product Poster Generation and Optimization

AAAI 2026technical

Product posters blend striking visuals with informative text to highlight the product and capture customer attention. However, crafting appealing posters and manually optimizing them based on online performance is laborious and resource-consuming. To address this, we introduce AutoPP, an automated p

Cited by 0SourcePDFScholar
2026

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

CVPR 2026

Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to real-wo

Cited by 0SourcecodeScholar
2026

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

CVPR 2026

Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image or text advertisements. However, their pipelines lack cross-modal perception and

Cited by 0SourcecodeScholar
2026

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

CVPR 2026

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. Within these frameworks, the policy update relies on importance-ratio clipping to constrain overconfident positive and negat

Cited by 0SourcecodeScholar
2026

ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation

AAAI 2026technical

Enabling multi-task adaptation in pre-trained Low-Rank Adaptation (LoRA) models is crucial for enhancing their generalization capabilities. Most existing pre-trained LoRA fusion methods decompose weight matrices, sharing similar parameters, while fusion divergent ones. However, this paradigm inevit

Cited by 0SourcePDFScholar
2026

InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation

CVPR 2026

E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed style. Recent diffusion models with fine-grained and efficient controllability have advanced product poster synthesis, yet t

Cited by 0SourceScholar
2026

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

AAAI 2026technical

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural g

Cited by 0SourcePDFScholar
2026

Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing

CVPR 2026

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models exhibit promising understanding capabilities, these strengt

Cited by 0SourceScholar
2026

RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers

AAAI 2026technical

The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource

Cited by 0SourcePDFScholar
2026

TR-DQ: Time-Rotation Diffusion Quantization

AAAI 2026technical

Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impa

Cited by 0SourcePDFScholar
2025

Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities

AAAI 2025technical

Text-to-Image generation (TTI) technologies are advancing rapidly, especially in the English language communities. However, apart from the user input language barrier problem, English-native TTI models inherently carry biases from their English world centric training data, which creates a dilemma fo…

2025

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

IJCAI 2025

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations wi

Cited by 0SourcePDFScholar
2025

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

ICCV 2025poster

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interact…

Cited by 0SourcePDFScholar
2025

Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts

ICCV 2025poster

Recently, learning-based stereo matching networks have advanced significantly.However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets.Leveraging Vision Foundation Models (VFMs)…

2025

Meta-Reinforcement Learning With Evolving Gradient Regularization

RA-L 2025

Deep reinforcement learning (DRL) typically requires reinitializing training for new tasks, limiting its generalization due to isolated knowledge transfer. Meta-reinforcement learning (Meta-RL) addresses this by enabling rapid adaptation through prior task experiences, yet existing gradient-based me

Cited by 1SourceScholar
2025

Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities

ACL 2025finding

Multimodal Large Language Models (MLLMs) have showcased exceptional Chain-of-Thought (CoT) reasoning ability in complex textual inference tasks including causal reasoning. However, will these causalities remain straightforward when crucial hints hide in visual details? If not, what factors might inf…

Cited by 0SourcePDFScholar
2025

PT-T2I/V: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Image/Video-Task

ICLR 2025poster

The global self-attention mechanism in diffusion transformers involves redundant computation due to the sparse and redundant nature of visual information, and the attention map of tokens within a spatial window shows significant similarity. To address this redundancy, we propose the Proxy-Tokenized…

2025

PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models

ICCV 2025poster

In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images as shown in Figure 1. Unlike previous diffusion-based models that treat layout planning and layout-to-image as two separate models, PlanGen j…

Cited by 0SourcePDFScholar
2025

WISA: World simulator assistant for physics-aware text-to-video generation

NeurIPS 2025spotlight

Recent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey ph…

Cited by 0SourcecodeScholar
2024

HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation

NeurIPS 2024poster

The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases include object missing, inconsistent lighting, conflicting view angles, etc. To effec…

2024

Offline Meta-Reinforcement Learning with Evolving Gradient Agreement

IROS 2024poster

Meta-Reinforcement Learning (Meta-RL) is a machine learning paradigm aimed at learning reinforcement learning policies that can quickly adapt to unseen tasks with few-shot data. Nevertheless, applying Meta-RL to real-world applications faces challenges due to the cost of data acquisition. To address…

Cited by 0SourceScholar
2023

Res-Tuning: A Flexible and Efficient Tuning Paradigm via Unbinding Tuner from Backbone

NeurIPS 2023poster

Parameter-efficient tuning has become a trend in transferring large-scale foundation models to downstream applications. Existing methods typically embed some light-weight tuners into the backbone, where both the design and the learning of the tuners are highly dependent on the base model. This work…