← Search

Yifeng Geng

12 accepted papers

2025

MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis

ICLR 2025poster

MetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the…

Cited by 2SourcePDFScholar
2025

UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

ICCV 2025poster

This paper presents UniPortrait, an innovative human image personalization framework that unifies single- and multi-ID customization with high face fidelity, extensive facial editability, free-form input description, and diverse layout generation. UniPortrait consists of only two plug-and-play modul…

2024

AnyText: Multilingual Visual Text Generation and Editing

ICLR 2024spotlight

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated im…

2024

Prune and Repaint: Content-Aware Image Retargeting for any Ratio

NeurIPS 2024poster

Image retargeting is the task of adjusting the aspect ratio of images to suit different display devices or presentation environments. However, existing retargeting methods often struggle to balance the preservation of key semantics and image quality, resulting in either deformation or loss of import…

2024

ShoeModel: Learning to Wear on the User-specified Shoes via Diffusion Model

ECCV 2024poster

"With the development of the large-scale diffusion model, Artificial Intelligence Generated Content (AIGC) techniques are popular recently. However, how to truly make it serve our daily lives remains an open question. To this end, in this paper, we focus on employing AIGC techniques in one filed of…

Cited by 2SourcePDFScholar
2023

DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving

IJCAI 2023poster

In the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO…

2023

FastInst: A Simple Query-Based Model for Real-Time Instance Segmentation

CVPR 2023poster

Recent attention in instance segmentation has focused on query-based models. Despite being non-maximum suppression (NMS)-free and end-to-end, the superiority of these models on high-accuracy real-time benchmarks has not been well demonstrated. In this paper, we show the strong potential of query-bas…

2023

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

IJCAI 2023poster

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a nove…

2023

Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception

ICASSP 2023accepted

Streaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns…

Cited by 0SourceScholar
2023

Optimal Proposal Learning for Deployable End-to-End Pedestrian Detection

CVPR 2023poster

End-to-end pedestrian detection focuses on training a pedestrian detection model via discarding the Non-Maximum Suppression (NMS) post-processing. Though a few methods have been explored, most of them still suffer from longer training time and more complex deployment, which cannot be deployed in the…

Cited by 19SourcePDFScholar
2023

Procontext: Exploring Progressive Context Transformer for Tracking

ICASSP 2023accepted

Existing Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with P…

Cited by 0SourceScholar
2023

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

ICCV 2023oral

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits…

Cited by 12PDFcodeScholar