OrientDream: Streamlining Text-to-3D Generation with Explicit Orientation Control
Yuzhong Huang, Fred Morstatter
Abstract
In the evolving landscape of text-to-3D technology, Dreamfusion [9] optimizes implicit representations like NeRF using Score Distillation Sampling (SDS) but faces limitations in both fidelity and speed. Specifically, it faces the multi-head Janus issue and exhibits a relatively slow optimization process. We present OrientDream, a camera orientation conditioned framework for efficient, multi-view consistent 3D generation from text prompts. OrientDream achieves this by pre-training a 2D text-to-image diffusion module with camera orientation features and utilizing data from MVImgNet. To shorten training time, we introduced a decoupled back-propagation technique, allowing for multiple updates of implicit parameters per optimization cycle. Our experiments reveal that our method not only produces high-quality NeRF models with consistent multi-view properties but also achieves an optimization speed significantly greater than existing methods, as quantified by comparative metrics.
BibTeX
@inproceedings{icassp2025_orientdreamstrea,
title = {OrientDream: Streamlining Text-to-3D Generation with Explicit Orientation Control},
author = {Yuzhong Huang and Fred Morstatter},
booktitle = {ICASSP 2025},
year = {2025}
}