← Search

Bojia Zi

10 accepted papers

2026

CTRL&SHIFT: High-quality Geometry-Aware Object Manipulation in Visual Generation

ICLR 2026poster

Object-level manipulation—relocating or reorienting objects in images or videos while preserving scene realism—is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core goals: background preservation, geometric consistency under viewpoi…

Cited by 0SourceScholar
2026

Refacade: Editing Object with Given Reference Texture

CVPR 2026

Recent advances in diffusion models have brought remarkable progress in image and video editing, yet some tasks remain underexplored. In this paper, we extend Object Retexture into video domain, which transfers local textures from a reference object to a target object in images or videos. To perform

Cited by 0SourcecodeScholar
2025

BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities

ICLR 2025poster

We introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same fr…

2025

CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

AAAI 2025technical

Video inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited numbe…

2025

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

NeurIPS 2025poster

Recent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated objects and visual artifacts. Furthermore, existing methods often rely on compu…

Cited by 0SourcecodeScholar
2025

Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

NeurIPS 2025poster

Video content editing has a wide range of applications. With the advancement of diffusion-based generative models, video editing techniques have made remarkable progress, yet they still remain far from practical usability. Existing inversion-based video editing methods are time-consuming and struggl…

Cited by 0SourcecodeScholar
2025

Taming Transformer Without Using Learning Rate Warmup

ICLR 2025poster

Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and an obviously lower learning rate, is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Tran…

Cited by 0SourcePDFScholar
2024

Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation

ECCV 2024poster

"Text-to-image generation has made significant advancements with the introduction of text-to-image diffusion models. These models typically consist of a language model that interprets user prompts and a vision model that generates corresponding images. As language and vision models continue to progr…

2021

Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better

ICCV 2021poster

Adversarial training is one effective approach for training robust deep neural networks against adversarial attacks. While being able to bring reliable robustness, adversarial training (AT) methods in general favor high capacity models, i.e., the larger the model the better the robustness. This tend…

Cited by 126PDFcodeScholar