← Search

Yuxuan Luo

11 accepted papers

2026

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation

CVPR 2026

Manual font design is an intricate process that transforms a stylistic visual concept into a coherent glyph set. This challenge persists in automated Few-shot Font Generation (FFG), where models struggle to preserve both structural integrity and stylistic fidelity from limited references. While auto

Cited by 0SourcecodeScholar
2026

FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing

AAAI 2026technical

Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies,

Cited by 0SourcePDFScholar
2026

Geometric Flow Grounding: A Unified Manifold Decoupling Framework for Dynamics Discovery and Verification

ICML 2026oral

Modeling complex dynamics from observational data is fundamental to scientific discovery and artificial intelligence. However, existing approaches ranging from Neural ODEs to diffusion models are often plagued by the entanglement of static state representations and instantaneous motion, leading to a…

Cited by 0SourceScholar
2025

CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model

ICCV 2025poster

Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-langu…

2025

DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance

ICCV 2025poster

While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coherence, which leads to their lower expressiveness and robustness. We propose a diffu…

Cited by 0SourcePDFScholar
2025

DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

AAAI 2025technical

Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetu…

Cited by 2SourcePDFScholar
2025

MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning

NeurIPS 2025poster

In this paper, we introduce knowledge image generation as a new task, alongside the Massive Multi-Discipline Multi-Tier Knowledge-Image Generation Benchmark (MMMG) to probe the reasoning capability of image generation models. Knowledge images have been central to human civilization and to the mechan…

Cited by 0SourceScholar
2024

CalliRewrite: Recovering Handwriting Behaviors from Calligraphy Images without Supervision

ICRA 2024poster

Human-like planning skills and dexterous manipulation have long posed challenges in the fields of robotics and artificial intelligence (AI). The task of reinterpreting calligraphy presents a formidable challenge, as it involves the decomposition of strokes and dexterous utensil control. Previous eff…

Cited by 0SourcecodeScholar
2024

Strike a Balance in Continual Panoptic Segmentation

ECCV 2024poster

"This study explores the emerging area of continual panoptic segmentation, highlighting three key balances. First, we introduce past-class backtrace distillation to balance the stability of existing knowledge with the adaptability to new information. This technique retraces the features associated w…

2023

Saving 100x Storage: Prototype Replay for Reconstructing Training Sample Distribution in Class-Incremental Semantic Segmentation

NeurIPS 2023poster

Existing class-incremental semantic segmentation (CISS) methods mainly tackle catastrophic forgetting and background shift, but often overlook another crucial issue. In CISS, each step focuses on different foreground classes, and the training set for a single step only includes images containing pix…

2017

WordSup: Exploiting Word Annotations for Character Based Text Detection

ICCV 2017poster

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and conven…

Cited by 250PDFScholar