← Search

Xuansong Xie

44 accepted papers

2025

MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis

ICLR 2025poster

MetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the…

Cited by 2SourcePDFScholar
2025

VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

AAAI 2025technical

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos.…

2024

3DToonify: Creating Your High-Fidelity 3D Stylized Avatar Easily from 2D Portrait Images

CVPR 2024poster

Visual content creation has aroused a surge of interest given its applications in mobile photography and AR/VR. Portrait style transfer and 3D recovery from monocular images as two representative tasks have so far evolved independently. In this paper we make a connection between the two and tackle t…

Cited by 2SourcePDFScholar
2024

AnyText: Multilingual Visual Text Generation and Editing

ICLR 2024spotlight

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated im…

2024

ChromaFusionNet (CFNet): Natural Fusion of Fine-Grained Color Editing

AAAI 2024technical

Digital image enhancement aims to deliver visually striking, pleasing images that align with human perception. While global techniques can elevate the image's overall aesthetics, fine-grained color enhancement can further boost visual appeal and expressiveness. However, colorists frequently face cha…

Cited by 1SourcePDFScholar
2024

DiffusionGAN3D: Boosting Text-guided 3D Generation and Domain Adaptation by Combining 3D GANs and Diffusion Priors

CVPR 2024poster

Text-guided domain adaptation and generation of 3D-aware portraits find many applications in various fields. However due to the lack of training data and the challenges in handling the high variety of geometry and appearance the existing methods for these tasks suffer from issues like inflexibility…

2024

DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation

ECCV 2024poster

"Text-to-3D generation, which synthesizes 3D assets according to an overall text description, has significantly progressed. However, a challenge arises when the specific appearances need customizing at designated viewpoints but referring solely to the overall description for generating 3D objects. F…

2024

En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic Data

CVPR 2024poster

We present En3D an enhanced generative scheme for sculpting high-quality 3D human avatars. Unlike previous works that rely on scarce 3D datasets or limited 2D collections with imbalanced viewing angles and imprecise pose priors our approach aims to develop a zero-shot 3D generative scheme capable of…

Cited by 10SourcePDFScholar
2024

Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction

AAAI 2024technical

Generative diffusion prior captured from the off-the-shelf denoising diffusion generative model has recently attained significant interest. However, several attempts have been made to adopt diffusion models to noisy inverse problems either fail to achieve satisfactory results or require a few thousa…

Cited by 4SourcePDFScholar
2024

InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data Pruning

ICLR 2024oral

Data pruning aims to obtain lossless performances with less overall cost. A common approach is to filter out samples that make less contribution to the training. This could lead to gradient expectation bias compared to the original data. To solve this problem, we propose InfoBatch, a novel framework…

2024

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

CVPR 2024highlight

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However there still remains a gap in providing fine-grained pixel-level…

2024

Pixel-Aware Stable Diffusion for Realistic Image Super-Resolution and Personalized Stylization

ECCV 2024poster

"Diffusion models have demonstrated impressive performance in various image generation, editing, enhancement and translation tasks. In particular, the pre-trained text-to-image stable diffusion models provide a potential solution to the challenging realistic image super-resolution (Real-ISR) and ima…

2024

ShoeModel: Learning to Wear on the User-specified Shoes via Diffusion Model

ECCV 2024poster

"With the development of the large-scale diffusion model, Artificial Intelligence Generated Content (AIGC) techniques are popular recently. However, how to truly make it serve our daily lives remains an open question. To this end, in this paper, we focus on employing AIGC techniques in one filed of…

Cited by 2SourcePDFScholar
2024

SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions

ECCV 2024poster

"Recent text-to-image generation methods such as ControlNet have achieved remarkable success in controlling image layouts, where the generated images by the default model are constrained to strictly follow the visual conditions (e.g., depth maps). However, in practice, the conditions usually provide…

2023

A Hierarchical Representation Network for Accurate and Detailed Face Reconstruction From In-the-Wild Images

CVPR 2023poster

Limited by the nature of the low-dimensional representational capacity of 3DMM, most of the 3DMM-based face reconstruction (FR) methods fail to recover high-frequency facial details, such as wrinkles, dimples, etc. Some attempt to solve the problem by introducing detail maps or non-linear operations…

2023

Boosting Novel Category Discovery Over Domains with Soft Contrastive Learning and All in One Classifier

ICCV 2023oral

Unsupervised domain adaptation (UDA) has proven to be highly effective in transferring knowledge from a label-rich source domain to a label-scarce target domain. However, the presence of additional novel categories in the target domain has led to the development of open-set domain adaptation (ODA) a…

Cited by 19PDFcodeScholar
2023

CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo

IJCAI 2023poster

The core of Multi-view Stereo(MVS) is the matching process among reference and source pixels. Cost aggregation plays a significant role in this process, while previous methods focus on handling it via CNNs. This may inherit the natural limitation of CNNs that fail to discriminate repetitive or incor…

Cited by 18SourcePDFScholar
2023

DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving

IJCAI 2023poster

In the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO…

2023

DDColor: Towards Photo-Realistic Image Colorization via Dual Decoders

ICCV 2023poster

Image colorization is a challenging problem due to multi-modal uncertainty and high ill-posedness. Directly training a deep neural network usually leads to incorrect semantic colors and low color richness. While transformer-based methods can deliver better results, they often rely on manually design…

Cited by 64PDFcodeScholar
2023

DamoFD: Digging into Backbone Design on Face Detection

ICLR 2023poster

Face detection (FD) has achieved remarkable success over the past few years, yet, these leaps often arrive when consuming enormous computation costs. Moreover, when considering a realistic situation, i.e., building a lightweight face detector under a computation-scarce scenario, such heavy computati…

2023

FastInst: A Simple Query-Based Model for Real-Time Instance Segmentation

CVPR 2023poster

Recent attention in instance segmentation has focused on query-based models. Despite being non-maximum suppression (NMS)-free and end-to-end, the superiority of these models on high-accuracy real-time benchmarks has not been well demonstrated. In this paper, we show the strong potential of query-bas…

2023

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

IJCAI 2023poster

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a nove…

2023

Improving Training and Inference of Face Recognition Models via Random Temperature Scaling

AAAI 2023technical

Data uncertainty is commonly observed in the images for face recognition (FR). However, deep learning algorithms often make predictions with high confidence even for uncertain or irrelevant inputs. Intuitively, FR algorithms can benefit from both the estimation of uncertainty and the detection of ou…

Cited by 10SourcePDFScholar
2023

Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception

ICASSP 2023accepted

Streaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns…

Cited by 0SourceScholar
2023

Optimal Proposal Learning for Deployable End-to-End Pedestrian Detection

CVPR 2023poster

End-to-end pedestrian detection focuses on training a pedestrian detection model via discarding the Non-Maximum Suppression (NMS) post-processing. Though a few methods have been explored, most of them still suffer from longer training time and more complex deployment, which cannot be deployed in the…

Cited by 19SourcePDFScholar
2023

PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-Modal Distillation and Super-Voxel Clustering

ICCV 2023poster

Semantic segmentation of point clouds usually requires exhausting efforts of human annotations, hence it attracts wide attention to a challenging topic of learning from unlabeled or weaker form of annotations. In this paper, we take the first attempt for fully unsupervised semantic segmentation of p…

Cited by 10PDFcodeScholar
2023

Procontext: Exploring Progressive Context Transformer for Tracking

ICASSP 2023accepted

Existing Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with P…

Cited by 0SourceScholar
2023

RSFNet: A White-Box Image Retouching Approach using Region-Specific Color Filters

ICCV 2023poster

Retouching images is an essential aspect of enhancing the visual appeal of photos. Although users often share common aesthetic preferences, their retouching methods may vary based on their individual preferences. Therefore, there is a need for white-box approaches that produce satisfying results and…

Cited by 14PDFcodeScholar
2023

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

ICCV 2023oral

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits…

Cited by 12PDFcodeScholar
2023

TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric Perspective

ICCV 2023poster

Vision Transformers (ViTs) have demonstrated powerful representation ability in various visual tasks thanks to their intrinsic data-hungry nature. However, we unexpectedly find that ViTs perform vulnerably when applied to face recognition (FR) scenarios with extremely large datasets. We investigate…

Cited by 34PDFcodeScholar
2022

ABPN: Adaptive Blend Pyramid Network for Real-Time Local Retouching of Ultra High-Resolution Photo

CVPR 2022poster

Photo retouching finds many applications in various fields. However, most existing methods are designed for global retouching and seldom pay attention to the local region, while the latter is actually much more tedious and time-consuming in photography pipelines. In this paper, we propose a novel ad…

Cited by 14PDFcodeScholar
2022

Active Boundary Loss for Semantic Segmentation

AAAI 2022technical

This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted b…

2022

Structure-Aware Flow Generation for Human Body Reshaping

CVPR 2022poster

Body reshaping is an important procedure in portrait photo retouching. Due to the complicated structure and multifarious appearance of human bodies, existing methods either fall back on the 3D domain via body morphable model or resort to keypoint-based image deformation, leading to inefficiency and…

Cited by 7PDFcodeScholar
2022

Unpaired Cartoon Image Synthesis via Gated Cycle Mapping

CVPR 2022poster

In this paper, we present a general-purpose solution to cartoon image synthesis with unpaired training data. In contrast to previous works learning pre-defined cartoon styles for specified usage scenarios (portrait or scene), we aim to train a common cartoon translator which can not only simultaneou…

Cited by 21PDFScholar
2021

EMLight: Lighting Estimation via Spherical Distribution Approximation

AAAI 2021technical

Illumination estimation from a single image is critical in 3D rendering and it has been investigated extensively in the computer vision and computer graphic research community. On the other hand, existing works estimate illumination by either regressing light parameters or generating illumination ma…

Cited by 165SourcePDFScholar
2021

Noise-Resistant Deep Metric Learning With Ranking-Based Instance Selection

CVPR 2021poster

The existence of noisy labels in real-world data negatively impacts the performance of deep learning models. Although much research effort has been devoted to improving robustness to noisy labels in classification tasks, the problem of noisy labels in deep metric learning (DML) remains open. In this…

Cited by 52PDFcodeScholar
2021

PPR10K: A Large-Scale Portrait Photo Retouching Dataset With Human-Region Mask and Group-Level Consistency

CVPR 2021poster

Different from general photo retouching tasks, portrait photo retouching (PPR), which aims to enhance the visual quality of a collection of flat-looking portrait photos, has its special and practical requirements such as human-region priority (HRP) and group-level consistency (GLC). HRP requires tha…

Cited by 57PDFcodeScholar
2021

Sparse Needlets for Lighting Estimation With Spherical Transport Loss

ICCV 2021poster

Accurate lighting estimation is challenging yet critical to many computer vision and computer graphics tasks such as high-dynamic-range (HDR) relighting. Existing approaches model lighting in either frequency domain or spatial domain which is insufficient to represent the complex lighting conditions…

Cited by 112PDFScholar
2021

Unbalanced Feature Transport for Exemplar-Based Image Translation

CVPR 2021poster

Despite the great success of GANs in images translation with different conditioned inputs such as semantic segmentation and edge map, generating high-fidelity images with reference styles from exemplars remains a grand challenge in conditional image-to-image translation. This paper presents a genera…

Cited by 235PDFScholar
2021

WaveFill: A Wavelet-Based Generation Network for Image Inpainting

ICCV 2021poster

Image inpainting aims to complete the missing or corrupted regions of images with realistic contents. The prevalent approaches adopt a hybrid objective of reconstruction and perceptual quality by using generative adversarial networks. However, the reconstruction loss and adversarial loss focus on sy…

Cited by 132PDFcodeScholar
2020

An AI-empowered Visual Storyline Generator

IJCAI 2020poster

Video editing is currently a highly skill- and time-intensive process. One of the most important tasks in video editing is to compose the visual storyline. This paper outlines Visual Storyline Generator (VSG), an artificial intelligence (AI)-empowered system that automatically generates visual story…

Cited by 0SourcePDFScholar
2020

Boosting Semantic Human Matting With Coarse Annotations

CVPR 2020oral

Semantic human matting aims to estimate the per-pixel opacity of the foreground human regions. It is quite challenging that usually requires user interactive trimaps and plenty of high quality annotated data. Annotating such kind of data is labor intensive and requires great skills beyond normal use…

Cited by 113PDFScholar
2019

Attention-Aware Multi-Stroke Style Transfer

CVPR 2019poster

Neural style transfer has drawn considerable attention from both academic and industrial field. Although visual effect and efficiency have been significantly improved, existing methods are unable to coordinate spatial distribution of visual attention between the content image and stylized image, or…

Cited by 215PDFcodeScholar