← Search

Guotao Liang

6 accepted papers

2026

Improved Masked Image Generation with Knowledge-Augmented Token Representations

AAAI 2026technical

Masked image generation (MIG) has demonstrated remarkable efficiency and high-fidelity images by enabling parallel token prediction. Existing methods typically rely solely on the model itself to learn semantic dependencies among visual token sequences. However, directly learning such semantic depend

Cited by 0SourcePDFScholar
2026

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

ICML 2026poster

Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence. However, this task remains challenging as it requires bridging discrete code representations with continuous visual dynamics. Existing optimization-…

Cited by 0SourceScholar
2025

AsyncDSB: Schedule-Asynchronous Diffusion Schrödinger Bridge for Image Inpainting

AAAI 2025technical

Image inpainting is an important image generation task, which aims to restore corrupted image from partial visible area. Recently, diffusion Schrödinger bridge methods effectively tackle this task by modeling the translation between corrupted and target images as a diffusion Schrödinger bridge proce…

Cited by 0SourcePDFScholar
2025

Empowering LLMs to Understand and Generate Complex Vector Graphics

CVPR 2025poster

The unprecedented advancements in Large Language Models (LLMs) have profoundly impacted natural language processing but have yet to fully embrace the realm of scalable vector graphics (SVG) generation. While LLMs encode partial knowledge of SVG data from web pages during training, recent findings su…

2025

Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text

CVPR 2025highlight

Image quantization is a crucial technique in image generation, aimed at learning a codebook that encodes an image into a discrete token sequence. Recent advancements have seen researchers exploring learning multi-modal codebook (i.e., text-aligned codebook) by utilizing image caption semantics, aimi…

Cited by 0SourcePDFScholar
2024

Codebook Transfer with Part-of-Speech for Vector-Quantized Image Modeling

CVPR 2024poster

Vector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook from scratch and in a code-independent manner to quantize con…

Cited by 11SourcePDFScholar