NeurIPS 2025poster0 citations

SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace Colorization

Junpyo Seo, HanbinKoo, Jieun Yook, Byung-Ro Moon

Abstract

We propose a novel diffusion-based framework for automatic colorization of Anime-style facial sketches, which preserves the structural fidelity of the input sketch while effectively transferring stylistic attributes from a reference image. Our approach builds upon recent continuous-time diffusion models, but departs from traditional methods that rely on predefined noise schedules, which often fail to maintain perceptual consistency across the generative trajectory. To address this, we introduce SSIMBaD (Sigma Scaling with SSIM-Guided Balanced Diffusion), a sigma-space transformation that ensures linear alignment of perceptual degradation, as measured by structural similarity. This perceptual scaling enforces uniform visual difficulty across timesteps, enabling more balanced and faithful reconstructions. On a large-scale Anime face dataset, SSIMBaD attains state-of-the-art structural fidelity and strong perceptual quality, with robust generalization to diverse styles and structural variations.

Diffusion ModelsNoise SchedulingPerceptual ConsistencyReference-guided GenerationConditional Diffusion ModelsGenerative ModelingSketch-to-Image TranslationStructural Similarity Index
BibTeX
@inproceedings{
seo2025ssimbad,
title={{SSIMB}aD: Sigma Scaling with {SSIM}-Guided Balanced Diffusion for AnimeFace Colorization},
author={Junpyo Seo and HanbinKoo and Jieun Yook and Byung-Ro Moon},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=qO7j8ymv5I}
}
SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace Colorization · NeurIPS 2025