← Search

Zhouhui Lian

27 accepted papers

2026

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation

CVPR 2026

Manual font design is an intricate process that transforms a stylistic visual concept into a coherent glyph set. This challenge persists in automated Few-shot Font Generation (FFG), where models struggle to preserve both structural integrity and stylistic fidelity from limited references. While auto

Cited by 0SourcecodeScholar
2026

IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments

AAAI 2026technical

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor Unmanned Aerial Vehicles (UAVs), indoor UAV-based VLN remains unde

Cited by 0SourcePDFScholar
2025

ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation

CVPR 2025poster

Multi-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative models. In this paper, we introduce the Anonymous Region Transformer (ART), which facilitates the direct generation of variab…

Cited by 4SourcePDFScholar
2025

CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model

ICCV 2025poster

Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-langu…

2025

Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting

CVPR 2025highlight

Personalized 3D avatar editing holds significant promise due to its user-friendliness and availability to applications such as AR/VR and virtual try-ons. Previous studies have explored the feasibility of 3D editing, but often struggle to generate visually pleasing results, possibly due to the unstab…

Cited by 0SourcePDFScholar
2025

MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning

NeurIPS 2025poster

In this paper, we introduce knowledge image generation as a new task, alongside the Massive Multi-Discipline Multi-Tier Knowledge-Image Generation Benchmark (MMMG) to probe the reasoning capability of image generation models. Knowledge images have been central to human civilization and to the mechan…

Cited by 0SourceScholar
2025

TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting

CVPR 2025poster

Physically Based Rendering (PBR) materials play a crucial role in modern graphics, enabling photorealistic rendering across diverse environment maps. Developing an effective and efficient algorithm that is capable of automatically generating high-quality PBR materials rather than RGB texture for 3D…

2024

3DToonify: Creating Your High-Fidelity 3D Stylized Avatar Easily from 2D Portrait Images

CVPR 2024poster

Visual content creation has aroused a surge of interest given its applications in mobile photography and AR/VR. Portrait style transfer and 3D recovery from monocular images as two representative tasks have so far evolved independently. In this paper we make a connection between the two and tackle t…

Cited by 2SourcePDFScholar
2024

CalliRewrite: Recovering Handwriting Behaviors from Calligraphy Images without Supervision

ICRA 2024poster

Human-like planning skills and dexterous manipulation have long posed challenges in the fields of robotics and artificial intelligence (AI). The task of reinterpreting calligraphy presents a formidable challenge, as it involves the decomposition of strokes and dexterous utensil control. Previous eff…

Cited by 0SourcecodeScholar
2024

DeepCalliFont: Few-Shot Chinese Calligraphy Font Synthesis by Integrating Dual-Modality Generative Models

AAAI 2024technical

Few-shot font generation, especially for Chinese calligraphy fonts, is a challenging and ongoing problem. With the help of prior knowledge that is mainly based on glyph consistency assumptions, some recently proposed methods can synthesize high-quality Chinese glyph images. However, glyphs in callig…

2024

En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic Data

CVPR 2024poster

We present En3D an enhanced generative scheme for sculpting high-quality 3D human avatars. Unlike previous works that rely on scarce 3D datasets or limited 2D collections with imbalanced viewing angles and imprecise pose priors our approach aims to develop a zero-shot 3D generative scheme capable of…

Cited by 10SourcePDFScholar
2024

TextNeRF: A Novel Scene-Text Image Synthesis Method based on Neural Radiance Fields

CVPR 2024poster

Acquiring large-scale well-annotated datasets is essential for training robust scene text detectors yet the process is often resource-intensive and time-consuming. While some efforts have been made to explore the synthesis of scene text images a notable gap remains between synthetic and authentic da…

2023

DeepVecFont-v2: Exploiting Transformers To Synthesize Vector Fonts With Higher Quality

CVPR 2023poster

Vector font synthesis is a challenging and ongoing problem in the fields of Computer Vision and Computer Graphics. The recently-proposed DeepVecFont achieved state-of-the-art performance by exploiting information of both the image and sequence modalities of vector fonts. However, it has limited capa…

2023

VecFontSDF: Learning To Reconstruct and Synthesize High-Quality Vector Fonts via Signed Distance Functions

CVPR 2023poster

Font design is of vital importance in the digital content design and modern printing industry. Developing algorithms capable of automatically synthesizing vector fonts can significantly facilitate the font design process. However, existing methods mainly concentrate on raster image generation, and o…

2022

Aesthetic Text Logo Synthesis via Content-Aware Layout Inferring

CVPR 2022poster

Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, few attention has been paid to this task which needs to take many factors (e.g., fonts, linguistics, topics, etc.) into cons…

Cited by 33PDFcodeScholar
2022

Unpaired Cartoon Image Synthesis via Gated Cycle Mapping

CVPR 2022poster

In this paper, we present a general-purpose solution to cartoon image synthesis with unpaired training data. In contrast to previous works learning pre-defined cartoon styles for specified usage scenarios (portrait or scene), we aim to train a common cartoon translator which can not only simultaneou…

Cited by 21PDFScholar
2021

CentripetalText: An Efficient Text Instance Representation for Scene Text Detection

NeurIPS 2021poster

Scene text detection remains a grand challenge due to the variation in text curvatures, orientations, and aspect ratios. One of the hardest problems in this task is how to represent text instances of arbitrary shapes. Although many methods have been proposed to model irregular texts in a flexible ma…

2020

Controllable Person Image Synthesis With Attribute-Decomposed GAN

CVPR 2020oral

This paper introduces the Attribute-Decomposed GAN, a novel generative model for controllable person image synthesis, which can produce realistic person images with desired human attributes (e.g., pose, head, upper clothes and pants) provided in various source inputs. The core idea of the proposed m…

Cited by 309PDFScholar
2017

Incremental Kernel Null Space Discriminant Analysis for Novelty Detection

CVPR 2017poster

Novelty detection, which aims to determine whether a given data belongs to any category of training data or not, is considered to be an important and challenging problem in areas of Pattern Recognition, Machine Learning, etc. Recently, kernel null space method (KNDA) was reported to have state-of-th…

Cited by 61PDFScholar