2026
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
CVPR 2026
Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs.