Source-linked AI summary
DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior
Jingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang, Wen Liu, Zhenda Xie, Yebin Liu
TL;DR
General 3D generation is constrained by limited 3D data and imperfect view consistency. DreamCraft3D lifts a 2D reference through geometry sculpting and bootstrapped texture refinement, producing high-fidelity assets with coherent 360° renderings.
Problem
General 3D-object generation remains difficult because extensive 3D training data are lacking, while fixed 2D priors provide per-view plausibility rather than global 3D consistency.
Method
DreamCraft3D hierarchically lifts a 2D reference image into 3D, using a view-conditioned diffusion prior for geometry and alternating DreamBooth adaptation with scene optimization for texture.
Results
DreamCraft3D produces high-fidelity 3D assets with compelling texture details, intricate geometry, and multi-view consistency in 360° renderings.
Takeaways & Limitations
Bootstrapped Score Distillation enables increasingly detailed texture while maintaining view consistency through an evolving, scene-specific diffusion prior.
Takeaways & Limitations
View-conditioned diffusion models are inherently 2D and cannot ensure perfect view consistency; SDS and VSD likewise suffer appearance and semantic shifts.
Abstract
from arXiv · showhide
We present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistency issue that existing works encounter. To sculpt geometries that render coherently, we perform score distillation sampling via a view-dependent diffusion model. This 3D prior, alongside several training strategies, prioritizes the geometry consistency but compromises the texture fidelity. We further propose Bootstrapped Score Distillation to specifically boost the texture. We train a personalized diffusion model, Dreambooth, on the augmented renderings of the scene, imbuing it with 3D knowledge of the scene being optimized. The score distillation from this 3D-aware diffusion prior provides view-consistent guidance for the scene. Notably, through an alternating optimization of the diffusion prior and 3D scene representation, we achieve mutually reinforcing improvements: the optimized 3D scene aids in training the scene-specific diffusion model, which offers increasingly view-consistent guidance for 3D optimization. The optimization is thus bootstrapped and leads to substantial texture boosting. With tailored 3D priors throughout the hierarchical generation, DreamCraft3D generates coherent 3D objects with photorealistic renderings, advancing the state-of-the-art in 3D content generation. Code available at https://github.com/deepseek-ai/DreamCraft3D.
1 INTRODUCTION
DreamCraft3D lifts a high-quality 2D reference image into 3D through hierarchical geometry sculpting and texture boosting, targeting detailed assets with holistic consistency.
- 3D content generation remains difficult because general 3D objects lack extensive training data.
- DreamCraft3D decomposes generation into 2D reference creation, geometry sculpting, and texture boosting.The stages mirror a manual artistic workflow and apply specialized techniques to each step.
- A view-conditioned diffusion model supplies 3D guidance during geometry sculpting to promote coherent and detailed geometry.
- Bootstrapped Score Distillation alternates diffusion-prior and 3D-scene optimization, using improved multi-view renderings to enhance texture while preserving view consistency.
- DreamCraft3D produces intricate geometries and realistic textures rendered coherently in 360°.The authors report improved texture and complexity over optimization-based approaches and realistic 360° renderings compared with image-to-3D techniques.
2 RELATED WORK
Prior 3D-generation methods span generative models, 3D-aware image synthesis, and 2D-to-3D lifting, but remain limited in open-domain consistency and large-view quality.
- GANs, autoregressive models, and diffusion-based approaches have been explored for generating 3D assets conditioned on images or text.
- 3D-aware image-generation methods target novel-view rendering with some level of 3D consistency, often using monocular depth prediction.
- Some category-focused methods produce photorealistic renderings but fall short for large views.
- 2D-to-3D lifting methods use CLIP or pretrained diffusion models to guide 3D representations and improve texture realism.
3 PRELIMINARIES
SDS and related distillation methods use pretrained diffusion models to optimize rendered 3D scenes, but fixed 2D priors provide limited global 3D consistency.
- SDS optimizes a rendered 3D representation by matching predicted and added noise under a pretrained text-to-image diffusion prior.
- Classifier-free guidance steers diffusion sampling toward conditional outputs, but large guidance weights can cause over-saturation and over-smoothing.
- VSD matches scores of noisy real and rendered images to obtain high-fidelity textures while modeling solutions as a distribution.
- SDS and VSD distill from fixed 2D distributions, which assure per-view plausibility rather than global 3D consistency and can cause appearance and semantic shifts.
- DreamCraft3D’s pipeline uses a 2D-generated image, view-conditioned diffusion guidance, and cyclic texture optimization with an evolving prior.
4 DREAMCRAFT3D
DreamCraft3D hierarchically lifts a high-quality 2D reference into coherent 3D geometry and detailed textures using complementary diffusion priors and staged optimization. Geometry consistency is prioritized first, then texture quality is improved through personalized, bootstrapped diffusion guidance.
- Hierarchical Generation: The pipeline generates a high-quality 2D image, then lifts it into 3D through cascaded geometry sculpting and texture refinement stages.This hierarchical design follows a coarse-to-fine artistic process.
- Geometry Sculpting: View-conditioned Zero-1-to-3 provides a 3D-aware prior for extrapolating plausible novel views from the reference image.Its distilled guidance is used for a 3D-aware SDS loss during geometry optimization.
- Geometry Sculpting: A hybrid SDS loss combines 2D and 3D diffusion priors because the 3D-aware model improves consistency but can impair generation quality when used alone.The hybrid formulation emphasizes the 3D diffusion prior with µ = 2, while DeepFloyd IF supplies coarse-geometry guidance for the 2D SDS term.
- Geometry Sculpting: Progressive view training and diffusion timestep annealing extend geometry from established views toward 360° while moving from global structure to finer details.The timestep range changes from [0.7, 0.85] early in optimization to [0.2, 0.5] later.
- Texture Boosting: Texture refinement switches to high-resolution Stable Diffusion gradients and optimizes mesh texture with the tetrahedral grid fixed, but geometry-stage priors leave textures blurry.DreamBooth is then trained on multi-view renderings to provide a scene-specific 3D prior for texture refinement.
- Texture Boosting: Bootstrapped Score Distillation alternates optimization of the 3D scene and its personalized diffusion prior, producing increasingly consistent texture guidance.As the mesh improves, reduced diffusion noise yields more consistent training renderings; two alternations suffice for consistent textures with rich details.
5 EXPERIMENTS
Experiments evaluate DreamCraft3D against text-to-3D and image-to-3D baselines using qualitative, quantitative, user-study, and ablation analyses. Results report stronger texture fidelity, view consistency, geometry quality, and user preference, while bootstrapping progressively improves renderings.
- Baselines: DreamCraft3D is compared with five baselines spanning text-to-3D and image-to-3D generation.The baselines are DreamFusion, Magic3D, ProlificDreamer, Make-it-3D, and Magic123.
- Qualitative comparison: Qualitatively, DreamCraft3D produces sharper geometry and texture details, including rich novel-view textures without multi-face Janus problems.
- Datasets: The benchmark contains 300 real or generated images, each paired with an alpha mask, predicted depth map, and text prompt.
- Quantitative comparison: Quantitatively, the method significantly surpasses baselines in texture consistency and fidelity across LPIPS, PSNR, Contextual Distance, and CLIP score.LPIPS and PSNR measure reference-view fidelity, Contextual Distance measures pixel-level congruence, and CLIP estimates semantic coherence.
- User study: 32 participants provided 480 responses, and users preferred DreamCraft3D in 92% of comparisons against alternative models.The study used 15 distinct prompt-and-image pairs and free-view rendering videos.
- Ablation study: Ablations show that the 3D prior supports globally consistent geometry, while BSD balances texture realism and consistency better than SDS or VSD.SDS produces smooth, over-saturated novel-view textures; VSD produces realistic but inconsistent textures.
- Stage-wise visualization: Across the hierarchical stages, geometry refinement improves high-resolution details with negligible texture change, whereas BSD significantly improves texture quality.DreamBooth bootstrapping further evolves renderings toward greater consistency and photorealism as the textured mesh is optimized.
6 CONCLUSION
DreamCraft3D combines geometry sculpting with Bootstrapped Score Distillation to produce high-fidelity, texture-rich, and multi-view-consistent 3D assets.
- DreamCraft3D introduces meticulous geometry sculpting and Bootstrapped Score Distillation for complex 3D asset generation.The approach targets plausible, coherent geometry and improved texture quality and consistency.
- Bootstrapped Score Distillation distills from an optimizing 3D-aware diffusion prior adapted to multi-view renderings of the instance.This improves texture quality and consistency during optimization.
- The method produces high-fidelity 3D assets with compelling texture details and multi-view consistency.
A.1 IMPLEMENTATION DETAILS
The implementation alternates scene-specific diffusion training with 3D optimization, using augmented multi-view renderings, geometric guidance, and staged sampling choices.
- Bootstrapped Score Distillation iteratively renders meshes, augments renderings with Gaussian noise, fine-tunes DreamBooth, and updates the 3D structure.The algorithm alternates diffusion-model and mesh optimization within a loop.
- A control net incorporates geometric normal information while inpainting unseen image segments to enforce view consistency.The visible reference-view region remains invariant while the inpainting model fills remaining segments.
- The Neus implementation uses a single-layer 32-unit MLP and density-based pruning every 10 iterations within an octree.The MLP predicts RGB color, volume density, and normal from multiresolution hash-encoded features.
- Camera and light sampling uses random augmentations, while material augmentation is frozen because it harmed training convergence.The point-light angular distance is sampled from ϕcam ∼U(0, π/3), with random point-light distance rcam ∼(7.5, 10).
- Geometry sculpting begins with time steps t ∼U(0.7, 0.85) and anneals to t ∼U(0.2, 0.50).Later iterations use t ∼U(0.2, 0.50).
A.2 ADDITIONAL EXPERIMENTS
Additional experiments show that DreamCraft3D can generate diverse, high-quality models from a single text prompt.
- DreamCraft3D generates an array of diverse models from a single text prompt, with remarkable quality.The method first translates the prompt into a reference image through 2D diffusion before image-based 3D creation.
A.3 LIMITATIONS
Additional qualitative results show photorealistic assets with compelling textures and improved 3D consistency, alongside a documented geometry-to-texture failure mode.
- The method occasionally incorporates frontal-view geometric details into texture because of depth ambiguity and inaccuracies in the depth prior.This failure is depicted in Figure 9.
- The method does not expressly separate material and lighting from the 2D reference image, leaving this aspect for future exploration.
- DreamCraft3D produces photorealistic 3D assets with compelling textural details and significantly improved 3D consistency.Figures 10–13 provide additional generated results.