← Search

Bonan Li

13 accepted papers

2026

FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation

AAAI 2026technical

Layout-to-Image generation has significantly advanced content creation by enabling the rendering of visual text under predefined spatial layouts. Current approaches achieve training-free layout guidance by constructing attention-based energy functions to derive correction gradients. In this paper, w

Cited by 0SourcePDFScholar
2025

A-PSRO: A Unified Strategy Learning Method with Advantage Metric for Normal-form Games

ICML 2025poster

Solving the Nash equilibrium in normal-form games with large-scale strategy spaces presents significant challenges. Open-ended learning frameworks, such as PSRO and its variants, have emerged as effective solutions. However, these methods often lack an efficient metric for evaluating strategy improv…

Cited by 0SourcePDFScholar
2025

CamPoint: Boosting Point Cloud Segmentation with Virtual Camera

CVPR 2025poster

Local features aggregation and global information perception are the fundamental to point cloud segmentation. However, existing works often fall short in effectively identifying semantic relevant neighbors and face challenges in endowing each point with high-level information. Here, we propose CamPo…

Cited by 0SourcePDFScholar
2025

CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creation

CVPR 2025highlight

Generating dense multiview images from text prompts is crucial for creating high-fidelity 3D assets. Nevertheless, existing methods struggle with space-view correspondences, resulting in sparse and low-quality outputs. In this paper, we introduce CoSER, a novel consistent dense Multiview Text-to-Ima…

2025

Control and Realism: Best of Both Worlds in Layout-to-Image without Training

ICML 2025poster

Layout-to-Image generation aims to create complex scenes with precise control over the placement and arrangement of subjects. Existing works have demonstrated that pre-trained Text-to-Image diffusion models can achieve this goal without training on any specific data; however, they often face challen…

Cited by 0SourcePDFScholar
2025

DreamHA: Towards High-Quality Human Animation with Image-to-Video Diffusion Models

ICASSP 2025accepted

Recent diffusion models have made significant advancements in generating lifelike videos from driving signals, including a reference character and a skeleton sequence. Nevertheless, these models often struggle with maintaining fidelity, as the generated results frequently deviate in character featur…

Cited by 0SourceScholar
2025

Feature out! Let Raw Image as Your Condition for Blind Face Restoration

ICML 2025poster

Blind face restoration (BFR), which involves converting low-quality (LQ) images into high-quality (HQ) images, remains challenging due to complex and unknown degradations. While previous diffusion-based methods utilize feature extractors from LQ images as guidance, using raw LQ images directly…

Cited by 0SourcePDFScholar
2025

MIRROR: Make Your Object-Level Multi-View Generation More Consistent with Training-Free Rectification

ICML 2025poster

Multi-view Diffusion has greatly advanced the development of 3D content creation by generating multiple images from distinct views, achieving remarkable photorealistic results. However, existing works are still vulnerable to inconsistent 3D geometric structures (commonly known as Janus Problem) and…

Cited by 0SourcePDFScholar
2025

StyO: Stylize Your Face in Only One-Shot

AAAI 2025technical

This paper focuses on face stylization with a single artistic target. Existing works for this task often fail to retain the source content while achieving geometry variation. Here, we present a novel StyO model, i.e., Stylize the face in only One-shot, to solve the above problem. In particular, StyO…

Cited by 9SourcePDFScholar
2024

Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis

CVPR 2024poster

Recent works in implicit representations such as Neural Radiance Fields (NeRF) have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters since the lack of explicit geometric constraints pose…

2023

Towards Consistent Video Editing with Text-to-Image Diffusion Models

NeurIPS 2023poster

Existing works have advanced Text-to-Image (TTI) diffusion models for video editing in a one-shot learning manner. Despite their low requirements of data and computation, these methods might produce results of unsatisfied consistency with text prompt as well as temporal sequence, limiting their appl…

Cited by 32SourcePDFScholar
2022

Shrinking Temporal Attention in Transformers for Video Action Recognition

AAAI 2022technical

Spatiotemporal modeling in an unified architecture is key for video action recognition. This paper proposes a Shrinking Temporal Attention Transformer (STAT), which efficiently builts spatiotemporal attention maps considering the attenuation of spatial attention in short and long temporal sequences.…

Cited by 13SourcePDFScholar