CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
Recent advancements in text-to-image synthesis have been largely propelled by diffusion-based models, yet achieving precise alignment between text prompts and generated images remains a persistent challenge. We find that this difficulty arises primarily from the limitations of conventional diffusion