Foundation Encoders Are All You Need for Preference-Aware Personalization
Hyungjin Kim, Seokho Ahn, Young-Duk Seo
Abstract
Personalized image generation based on user behavior reflects individual preferences with minimal user intervention. However, existing studies often rely on inaccurate profiling, high resource costs, and model-specific designs, which jointly restrict creativity, diversity, and generality. To address these limitations, we propose FAN, a novel approach that enables preference-aware personalization using only foundation encoders, without additional structures. FAN performs tailored profiling to capture user preferences and reconstructs transformer-based encoders to integrate them while preserving target fidelity. Experiments show that FAN achieves robust personalization across various foundation text-to-image models. It also extends to applications such as CLIP retrieval, unCLIP, and vision-language models, seamlessly integrating into diverse encoders without fine-tuning.
BibTeX
@inproceedings{cvpr2026_foundationencode,
title = {Foundation Encoders Are All You Need for Preference-Aware Personalization},
author = {Hyungjin Kim and Seokho Ahn and Young-Duk Seo},
booktitle = {CVPR 2026},
year = {2026}
}