Source-linked AI summary
NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° Views
Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, Zhangyang Wang
TL;DR
Single-image 3D reconstruction is difficult because existing methods rely on tedious manual work, limited generalization, or accurate depth. NeuralLift-360 combines NeRF with CLIP-guided diffusion priors and scale-invariant depth ranking, and reports improved performance over current state-of-the-art methods.
Problem
Single-image 3D reconstruction must handle hidden content and unreliable depth when creating plausible 360° objects from in-the-wild photos.
Method
NeuralLift-360 combines a NeRF representation with CLIP-guided diffusion priors, single-image adaptation, and scale-invariant depth ranking supervision.
Results
NeuralLift-360 outperforms current state-of-the-art methods on real and synthetic images.
Takeaways & Limitations
The framework demonstrates plausible 3D objects with 360° views that correspond well to given in-the-wild reference images.
Takeaways & Limitations
The method renders at 128 × 128 resolution and does not include challenging cases with multiple occluding objects.
Abstract
from arXiv · showhide
Virtual reality and augmented reality (XR) bring increasing demand for 3D content. However, creating high-quality 3D content requires tedious work that a human expert must do. In this work, we study the challenging task of lifting a single image to a 3D object and, for the first time, demonstrate the ability to generate a plausible 3D object with 360° views that correspond well with the given reference image. By conditioning on the reference image, our model can fulfill the everlasting curiosity for synthesizing novel views of objects from images. Our technique sheds light on a promising direction of easing the workflows for 3D artists and XR designers. We propose a novel framework, dubbed NeuralLift-360, that utilizes a depth-aware neural radiance representation (NeRF) and learns to craft the scene guided by denoising diffusion models. By introducing a ranking loss, our NeuralLift-360 can be guided with rough depth estimation in the wild. We also adopt a CLIP-guided sampling strategy for the diffusion prior to provide coherent guidance. Extensive experiments demonstrate that our NeuralLift-360 significantly outperforms existing state-of-the-art baselines. Project page: https://vita-group.github.io/NeuralLift-360/
1. Introduction
NeuralLift-360 addresses single-image 3D object creation by combining NeRF, diffusion priors, CLIP guidance, and scale-invariant depth supervision to produce 360° views from in-the-wild photos.
- Manual 3D content creation remains tedious, while existing automatic pipelines commonly require hundreds of images and multi-view stereo.
- Single-image approaches suffer from poor generalization or dependence on accurate depth, limiting their performance on in-the-wild images.
- NeuralLift-360 converts diverse in-the-wild 2D photos into 3D content with 360° views using diffusion priors and monocular depth cues.
- The framework uses NeRF as its scene representation and integrates diffusion-model prior knowledge to lift a single image into 3D.
- CLIP-guided sampling combines diffusion prior knowledge with the reference image, while scale-invariant depth supervision uses ranking information.
2. Related Works
Related work infers 3D content from single images using depth-based reprojection, learned 3D priors, and NeRF-like representations, but these approaches face depth instability, artifacts, and limited single-view reconstruction.
- Depth-based methods back-project monocular depth into point clouds or multi-plane images, then inpaint holes in novel views.
- Depth estimators can be unstable and inpainting may introduce artifacts that make rendered images look fake.
- Other approaches train on large-scale 3D asset datasets and directly predict a 3D object from a single image.
- Unlike most NeRF-like models requiring several views, this work focuses on reconstructing 3D content from a single view.
3. Preliminaries
The preliminaries describe NeRF volumetric rendering and denoising diffusion models, while the pipeline overview connects these components through diffusion and rough-depth supervision.
- Neural Radiance Fields: NeRF samples 5D coordinates comprising 3D location and viewing direction along camera rays.
- Neural Radiance Fields: NeRF maps sampled coordinates to color and volume density, then uses volumetric rendering to accumulate color along each ray.
- NeuralLift-360 pipeline: The proposed pipeline derives a CLIP-guided diffusion prior loss and uses ranking loss to incorporate rough-depth information.
- Neural Radiance Fields: The implementation uses Instant-NGP with a multiresolution hash table to reduce computation when inferring color and density.
- Denoising Diffusion Models: Denoising diffusion models generate images by progressively transforming Gaussian noise into samples from a target data distribution.
4. Method
NeuralLift-360 reconstructs a 3D object from one image by combining NeRF with diffusion-model priors, while grounding the result in the reference appearance. It adapts diffusion guidance to in-the-wild images and uses relative depth supervision to handle imperfect geometry cues.
- Probabilistic-Driven 3D Lifting: NeuralLift-360 represents the scene with NeRF and uses diffusion priors to hallucinate unseen views while reconstructing a 3D object from one image.The training combines prior distillation from a generative model with a direct loss on the given image.
- Probabilistic-Driven 3D Lifting: Camera poses are sampled during training, and rendered views receive both image-referenced guidance and reference-free diffusion regularization.The referenced term encourages appearance consistency with the input, whereas the non-referenced term supplies prior-based hallucination.
- CLIP-guided Diffusion Priors: CLIP-guided mixed reference guidance injects similarity to the input image into diffusion sampling while retaining text-conditioned generative priors.The method measures reference compatibility in CLIP feature space and alternates fixed reference-view supervision with stochastic camera-view supervision.
- CLIP-guided Diffusion Priors: Unlike DreamFusion, NeuralLift-360 explicitly grounds the generated 3D scene in the given image to maintain consistent appearance with the specified object.
- Domain Adaption to In-the-wild Images: Single-image adaptation fine-tunes diffusion on the reference while using augmentations to preserve diverse outputs rather than reproducing one identical image.This diversity is needed because identical supervision across views can otherwise cause NeRF to produce an isotropic surface.
- Supervision from Rough Depth: Relative depth ranking supervision avoids relying on unreliable absolute depth scales and provides additional geometry supervision alongside photometric and diffusion losses.The ranking loss uses pairwise ordering from monocular depth estimates, whose absolute values are insufficient for recovering plausible 3D shapes in the wild.
5. Experiments
Experiments evaluate NeuralLift-360 on synthetic and real images, including quantitative comparisons and ablations of its supervision and CLIP-guided diffusion prior. The method achieves the best reported CLIP distance and produces visually pleasing, 3D-consistent novel views.
- Training recipe: The training recipe combines diverse camera sampling, foreground-aware diffusion guidance, geometry regularization, and timestep annealing.These components target coherent 3D geometry, foreground emphasis, artifact reduction, and progressively more accurate diffusion guidance.
- Experimental setup: Experiments use synthetic and real-world images, with foreground segmentation and monocular depth estimation applied to each image.Synthetic images are generated with PNDM; real images come from CO3Dv2 or Google searches.
- Quantitative comparison: The method achieves the best CLIP distance among existing approaches across eight evaluated scenes.Evaluation uses a different CLIP model from training, held-out viewing directions, and 100 rendered images per scene and method.
- Visual comparison: Visual comparisons show that NeuralLift-360 synthesizes visually pleasing novel views with 3D consistency.Compared methods are affected by inaccurate monocular depth or unconstrained geometry.
- Ablation study: The full model performs best in ablations, while removing ranking loss, diffusion guidance, or fine-tuning produces unreliable depth, blurry geometry, or incorrect objects.The no-fine-tuning variant introduces lettuce absent from the reference image.
- Ablation study: Removing CLIP guidance yields a strange shape that does not resemble the reference, whereas the full model produces a better-matching shape.This ablation directly evaluates CLIP guidance in the diffusion prior.
6. Conclusions and Future Work
NeuralLift-360 lifts in-the-wild 2D photos into 3D objects with 360-degree views using CLIP-guided diffusion priors and scale-invariant depth ranking. The authors report strong results but identify resolution and scene-complexity limits for future work.
- Conclusion: NeuralLift-360 lifts an in-the-wild 2D photo into a 3D object with 360-degree views.The framework combines probabilistic-driven 3D lifting with CLIP-guided diffusion priors and scale-invariant depth ranking loss.
- Limitations: The target resolution is 128 × 128, which remains behind large generative models.This is stated as a limitation of the current implementation.
- Future work: Scenes with multiple occluding objects are outside the method’s assumptions and require further exploration.The authors plan to expand the method to more general scenarios.
A. Introduction
The supplementary material introduces detailed derivations, additional training recipes, implementation details, and experimental comparisons for NeuralLift-360.
- The supplementary material covers the probabilistic-driven diffusion prior, additional training recipes, implementation details, and experimental comparisons.
B. Detailed Derivation.
This section formalizes the objective and notation for NeuralLift-360, including the image, text description, 3D scene representation, and rendering function.
- Notations: The derivation defines an image y, text description z, and 3D scene representation V, with h(V, Φ) rendering V from camera pose Φ.V may represent a radiance volume or an implicit neural representation.
B.1. Pose Conditioned Latent Model
The model formulates 3D reconstruction as posterior inference over a latent scene, camera pose, and rendered image, then replaces the prior-based objective with diffusion score matching and guidance.
- Pose-conditioned latent formulation: The renderer h(V, Φ) maps the 3D representation and camera pose into the image domain for likelihood and prior modeling.The likelihood and prior are rewritten through h(V, Φ), connecting 3D estimation to image-space supervision.
- Pose-conditioned latent formulation: The posterior decomposes into an image likelihood and a scene prior, with camera pose introduced as a latent variable.The derivation assumes camera pose Φ is independent of z before rewriting the likelihood.
- Diffusion evidence lower bound: Jensen’s inequality produces an evidence lower bound, which is then approximated with score matching using a pretrained generative prior.The score function is defined as ϵθ(x|y, z) = ∇log pθ(x|y, z).
- Diffusion guidance: Bayesian rewriting combines a reference-conditioned diffusion model with an off-the-shelf text-to-image diffusion model to compute the conditional score.The guidance formulation begins by expressing pθ(x|y, z) through Bayesian rule and then deriving its score function.
- Diffusion guidance: Classifier-free guidance further modifies the diffusion objective through a guidance-strength parameter ω.The supplied passage identifies ω as the classifier guidance strength.
B.3. Computational Specifications
The implementation combines reference-conditioned diffusion, pseudo-depth supervision, CLIP similarity guidance, and camera sampling that alternates between reference and novel views.
- Implementation components: The final loss combines implementation components derived from the probabilistic objective, including diffusion guidance and reference-based supervision.The section introduces several implementations leading toward the final loss function.
- Implementation components: CLIP guidance measures reference-image similarity with an inner product in CLIP embedding space, using a CLIP image encoder.The inner product is used as a similarity score, although it is not rigorously a metric.
- Camera sampling: Camera poses are sampled on a unit sphere with orientations directed toward the origin, with a Bernoulli switch controlling reference-view utilization.When the switch is off, the rendered image is modeled around the reference image with Gaussian noise.
- Implementation components: Pseudo-depth supervision regularizes the objective with a ranking loss between rendered depth and monocular depth estimated from the reference image.The renderer independently produces RGB and depth, while p(d|y) is defined from the ranking loss.
C. Additional Training Recipe
The training recipe improves geometry and appearance through surface-based shading, geometric regularization, depth smoothing, multiresolution normals, background modeling, and controlled diffusion guidance.
- Shading and geometry: The implementation represents geometry explicitly for shading and uses surface normals to improve geometry quality.Surface normals are defined as negative density gradients, and the implementation replaces view-dependent radiance with surface-based geometry.
- Shading and geometry: Diffuse reflectance combines surface color, point-light illumination, ambient light, and a randomly perturbed light position near the camera.The rendering equation uses the ray direction, surface normal, point-light color, and ambient-light color.
- Geometry regularization: Without geometry regularization, unobserved regions can become arbitrary, and shading makes image quality dependent on geometry quality.This motivates additional priors and regularization beyond image reconstruction alone.
- Geometry regularization: Backward-facing normal penalties address flat objects and foggy floaters caused by semi-transparent rear surfaces that reproduce the front view.The penalty targets surfaces whose normals face backward relative to the viewing direction.
- Geometry regularization: Distortion, sparsity, and inverse-depth smoothing losses constrain floating geometry, occupancy, and spiky depth predictions.Inverse-depth smoothing makes rendered depth consistent with rendered RGB rather than only preserving internal rankings.
- Rendering and background: Low-resolution normal rendering reduces the computational burden of finite-difference density gradients during patch-based training.The implementation queries eight neighboring 3D points and renders normals at 100 × 100 resolution.
- Rendering and background: A separate background NeRF generates view-dependent backgrounds, reducing the foreground NeRF’s burden and helping remove floating artifacts.The background module addresses the mismatch between diffusion-model training images and foreground-only reconstruction.
- Diffusion guidance: The method avoids the range mismatch associated with large diffusion guidance weights because its sampling setup differs from iterative text-to-image generation.The passage explains that the usual [−1, 1] range issue does not arise in the same way here.