← Search

Wangmeng Xiang

10 accepted papers

2025

MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis

ICLR 2025poster

MetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the…

Cited by 2SourcePDFScholar
2024

AnyText: Multilingual Visual Text Generation and Editing

ICLR 2024spotlight

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated im…

2023

A Benchmark for Chinese-English Scene Text Image Super-Resolution

ICCV 2023poster

Scene Text Image Super-resolution (STISR) aims to recover high-resolution (HR) scene text images with visually pleasant and readable text content from the given low-resolution (LR) input. Most existing works focus on recovering English texts, which have simple structures in the characters, while lit…

Cited by 12PDFcodeScholar
2023

DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving

IJCAI 2023poster

In the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO…

2023

Generative Action Description Prompts for Skeleton-based Action Recognition

ICCV 2023poster

Skeleton-based action recognition has recently received considerable attention. Current approaches to skeleton-based action recognition are typically formulated as one-hot classification tasks and do not fully exploit the semantic relations between actions. For example, "make victory sign" and "thum…

Cited by 65PDFcodeScholar
2023

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

IJCAI 2023poster

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a nove…

2023

MDQE: Mining Discriminative Query Embeddings To Segment Occluded Instances on Challenging Videos

CVPR 2023poster

While impressive progress has been achieved, video instance segmentation (VIS) methods with per-clip input often fail on challenging videos with occluded objects and crowded scenes. This is mainly because instance queries in these methods cannot encode well the discriminative embeddings of instances…

2023

Procontext: Exploring Progressive Context Transformer for Tracking

ICASSP 2023accepted

Existing Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with P…

Cited by 0SourceScholar
2022

Spatiotemporal Self-Attention Modeling with Temporal Patch Shift for Action Recognition

ECCV 2022poster

"Transformer-based methods have recently achieved great advancement on 2D image-based vision tasks. For 3D video-based tasks such as action recognition, however, directly applying spatiotemporal transformers on video data will bring heavy computation and memory burdens due to the largely increased n…

2021

Real-World Video Super-Resolution: A Benchmark Dataset and a Decomposition Based Learning Scheme

ICCV 2021poster

Video super-resolution (VSR) aims to improve the spatial resolution of low-resolution (LR) videos. Existing VSR methods are mostly trained and evaluated on synthetic datasets, where the LR videos are uniformly downsampled from their high-resolution (HR) counterparts by some simple operators (e.g., b…

Cited by 62PDFcodeScholar