← Search

Ye Ma

10 accepted papers

2026

Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

ICML 2026poster

Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equipped with large language model (LLM)-based text encoders, remain text-pixel mappers -- they employ LLMs merely as text enco…

Cited by 0SourceScholar
2025

D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching

AAAI 2025technical

Videos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying link, etc. Adding proper sound effects (SFX) to such moments, or video decorati…

Cited by 0SourcePDFScholar
2025

Improving Preference Alignment of LLM with Inference-Free Self-Refinement

EMNLP 2025

Large language models (LLMs) develop the in-context learning capability through pretraining and instruction tuning, enabling task adaptation without parameter updates. Self-refinement is a manifestation of this capability, which allows LLMs to iteratively refine the output using self-generated feedb

2025

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

ICML 2025poster

We introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the \textbf{AR} modeling principle. The continuous treatment of visual signals minimize…

Cited by 8SourcePDFScholar
2022

Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs

IJCAI 2022poster

In this paper, we study the graphic layout generation problem of producing high-quality visual-textual presentation designs for given images. We note that image compositions, which contain not only global semantics but also spatial information, would largely affect layout results. Hence, we propose…

2022

DVS-Voltmeter: Stochastic Process-Based Event Simulator for Dynamic Vision Sensors

ECCV 2022poster

"Recent advances in deep learning for event-driven applications with dynamic vision sensors (DVS) primarily rely on training over simulated data. However, most simulators ignore various physics-based characteristics of real DVS, such as the fidelity of event timestamps and comprehensive noise effect…