← Search

Xiaoshi Wu

9 accepted papers

2026

Latent Diffusion Model without Variational Autoencoder

ICLR 2026poster

Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with Variational Autoencoders (VAEs). While effective for high-fidelity synthesis, this VAE+Diffusion paradigm still suffers from limited training and inference efficiency, along with poor transferabil…

Cited by 0SourcecodeScholar
2024

Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation

ECCV 2024poster

"Video outpainting is a challenging task, aiming at generating video content outside the viewport of the input video while maintaining inter-frame and intra-frame consistency. Existing methods fall short in either generation quality or flexibility. We introduce (Mastering Video Outpainting Through I…

2024

CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

NeurIPS 2024poster

Diffusion models have demonstrated great success in the field of text-to-image generation. However, alleviating the misalignment between the text prompts and images is still challenging. We break down the problem into two causes: concept ignorance and concept mismapping. To tackle the two challenges…

2024

Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models

ECCV 2024poster

"Optimizing a text-to-image diffusion model with a given reward function is an important but underexplored research area. In this study, we propose Deep Reward Tuning (DRTune), an algorithm that directly supervises the final output image of a text-to-image diffusion model and back-propagates through…

Cited by 14SourcePDFScholar
2023

CORA: Adapting CLIP for Open-Vocabulary Detection With Region Prompting and Anchor Pre-Matching

CVPR 2023poster

Open-vocabulary detection (OVD) is an object detection task aiming at detecting objects from novel categories beyond the base categories on which the detector is trained. Recent OVD methods rely on large-scale visual-language pre-trained models, such as CLIP, for recognizing novel objects. We identi…

2023

Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

ICCV 2023poster

Recent years have witnessed a rapid growth of deep generative models, with text-to-image models gaining significant attention from the public. However, existing models often generate images that do not align well with human preferences, such as awkward combinations of limbs and facial expressions. T…

Cited by 122PDFcodeScholar
2022

Uni-Perceiver: Pre-Training Unified Architecture for Generic Perception for Zero-Shot and Few-Shot Tasks

CVPR 2022poster

Biological intelligence systems of animals perceive the world by integrating information in different modalities and processing simultaneously for various tasks. In contrast, current machine learning research follows a task-specific paradigm, leading to inefficient collaboration between tasks and hi…

Cited by 147PDFScholar
2021

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

ICCV 2021poster

The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world's landmarks from tourist photos. However, a major source of information available for these 3D-augmented c…

Cited by 20PDFcodeScholar