← Search

Mushui Liu

11 accepted papers

2026

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

ICML 2026poster

Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome-based reward to all preceding denoising steps without distinguishing the local effect of each step. Moreover, current group-wise ranking mainly compare…

Cited by 0SourceScholar
2026

CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation

AAAI 2026technical

While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-image comprehension tasks. This limitation stems from their predominant reliance on

Cited by 0SourcePDFScholar
2026

DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation

CVPR 2026

The unified autoregressive (AR) model excels at multimodal understanding and generation. However, its full potential in the domain of customized image generation has yet to be fully realized.Existing customization approaches for unified AR models face a fundamental dilemma: adaptation-based methods

Cited by 0SourcecodeScholar
2026

FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation

AAAI 2026technical

Recent unified models have demonstrated that the reasoning capacity of Multimodal Large Language Models (MLLMs) can be leveraged to facilitate diffusion-based image generation with impressive flexibility and performance. However, approaches that rely heavily on MLLMs for high-level semantic encoding

Cited by 0SourcePDFScholar
2026

MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement

ICLR 2026poster

Multi-subject personalized generation presents unique challenges in maintaining identity fidelity and semantic coherence when synthesizing images conditioned on multiple reference subjects. Existing methods often suffer from identity blending and attribute leakage due to inadequate modeling of how d…

Cited by 0SourcecodeScholar
2025

Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning

AAAI 2025technical

Few-shot learning (FSL) aims to recognize new concepts using a limited number of visual samples. Existing methods attempt to incorporate semantic information into the limited visual data for category understanding. However, these methods often enrich class-level feature representations with abstract…

Cited by 9SourcePDFScholar
2025

Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition

AAAI 2025technical

In this paper, we propose a novel Temporal Sequence-Aware-Model (TSAM) for few-shot action recognition (FSAR), which incorporates a sequential perceiver adapter into the pre-training framework, to integrate both the spatial information and the sequential temporal dynamics into the feature embeddings…

Cited by 6SourcePDFScholar
2025

LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation

AAAI 2025technical

Diffusion models have exhibited substantial success in text-to-image generation. However, they often encounter challenges when dealing with complex and dense prompts involving multiple objects, attribute binding, and long descriptions. In this paper, we propose a novel framework called LLM4GEN, whic…

Cited by 20SourcePDFScholar
2025

MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

AAAI 2025technical

Auto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I g…

2025

TFCustom: Customized Image Generation with Time-Aware Frequency Feature Guidance

CVPR 2025highlight

Subject-driven image personalization has seen notable advancements, especially with the advent of the ReferenceNet paradigm. ReferenceNet excels in integrating image reference features, making it highly applicable in creative and commercial settings. However, current implementations of ReferenceNet…

Cited by 0SourcePDFScholar
2024

Improving Zero-Shot Generalization for CLIP with Variational Adapter

ECCV 2024poster

"The excellent generalization capability of pre-trained Vision-Language Models (VLMs) makes fine-tuning VLMs for downstream zero-shot tasks a popular choice. Despite achieving promising performance in the professionality of base classes, most existing fine-tuned methods suffer from feature confusion…

Cited by 8SourcePDFScholar