ICLR 2024poster18 citations

TOSS: High-quality Text-guided Novel View Synthesis from a Single Image

Yukai Shi, Jianan Wang, He CAO, Boshi Tang, Xianbiao Qi, Tianyu Yang, Yukun Huang, Shilong Liu

Abstract

In this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image. While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often result in implausible NVS generations. To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space. TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details. Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with higher-quality NVS results and faster convergence. We further support these results with comprehensive ablations that underscore the effectiveness and potential of the introduced semantic guidance and architecture design.

novel view synthesis3D object synthesisdiffusion model
BibTeX
@inproceedings{
shi2024toss,
title={{TOSS}: High-quality Text-guided Novel View Synthesis from a Single Image},
author={Yukai Shi and Jianan Wang and He CAO and Boshi Tang and Xianbiao Qi and Tianyu Yang and Yukun Huang and Shilong Liu and Lei Zhang and Heung-Yeung Shum},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024},
url={https://openreview.net/forum?id=9ZUYJpvIys}
}
TOSS: High-quality Text-guided Novel View Synthesis from a Single Image · ICLR 2024