← Search

Yuhang Ma

10 accepted papers

2026

NAMI: Efficient Image Generation via Bridged Progressive Rectified Flow Transformers

CVPR 2026

Flow-based Transformer models have achieved state-of-the-art image generation performance, but often suffer from high inference latency and computational cost due to their large parameter sizes. To improve inference efficiency without compromising quality, we propose Bridged Progressive Rectified Fl

Cited by 0SourceScholar
2026

RefTon: Reference person shot assist virtual Try-on

CVPR 2026

We introduce RefTon, a flux-based person-to-person virtual try-on framework that enhances garment realism through unpaired visual references. Unlike conventional approaches that rely on complex auxiliary inputs such as body parsing and warped mask or require finely designed extract branches to proce

Cited by 0SourcecodeScholar
2025

Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities

AAAI 2025technical

Text-to-Image generation (TTI) technologies are advancing rapidly, especially in the English language communities. However, apart from the user input language barrier problem, English-native TTI models inherently carry biases from their English world centric training data, which creates a dilemma fo…

2025

Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision Synthesis

IJCAI 2025

Multimodal perception, which integrates vision and touch, is increasingly demonstrating its significance in domains such as embodied intelligence and human-computer interaction. However, in open-world scenarios, multimodal data streams face significant challenges, including catastrophic forgetting a

2025

LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation

AAAI 2025technical

Diffusion models have exhibited substantial success in text-to-image generation. However, they often encounter challenges when dealing with complex and dense prompts involving multiple objects, attribute binding, and long descriptions. In this paper, we propose a novel framework called LLM4GEN, whic…

Cited by 20SourcePDFScholar
2025

PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models

ICCV 2025poster

In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images as shown in Figure 1. Unlike previous diffusion-based models that treat layout planning and layout-to-image as two separate models, PlanGen j…

Cited by 0SourcePDFScholar
2025

Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection

AAAI 2025technical

Recent advances in text-to-image diffusion models have spurred significant interest in continuous story image generation. In this paper, we introduce Storynizor, a model capable of generating coherent stories with strong inter-frame character consistency, effective foreground-background separation,…

Cited by 1SourcePDFScholar
2022

Continual Federated Learning Based on Knowledge Distillation

IJCAI 2022poster

Federated learning (FL) is a promising approach for learning a shared global model on decentralized data owned by multiple clients without exposing their privacy. In real-world scenarios, data accumulated at the client-side varies in distribution over time. As a consequence, the global model tends t…

Cited by 86SourcePDFScholar
2021

Beelines: Motion Prediction Metrics for Self-Driving Safety and Comfort

ICRA 2021poster

The commonly used metrics for motion prediction do not correlate well with a self-driving vehicle’s system-level performance. The most common metrics are average displacement error (ADE) and final displacement error (FDE), which omit many features, making them poor self-driving performance indicator…

Cited by 3SourceScholar