Source-linked AI summary
LERF: Language Embedded Radiance Fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, Matthew Tancik
TL;DR
NeRFs represent scenes photorealistically but lack semantic meaning, motivating language-based 3D interaction. LERF distills off-the-shelf CLIP embeddings into a dense, multi-scale NeRF field, yielding broad, real-time, 3D-consistent language relevancy maps. Its scope is bounded by CLIP and NeRF limitations and by the need for calibrated, high-quality multi-view captures.
Problem
NeRFs produce colorful density fields without meaning, while natural-language 3D interaction requires multi-scale semantics spanning long-tail and abstract concepts.
Method
LERF jointly optimizes a position-and-scale-conditioned language field inside NeRF using multi-scale CLIP embeddings from training-view image crops.
Results
LERF supports broad natural-language queries with localized, 3D-consistent relevancy maps in real time across diverse in-the-wild scenes.
Takeaways & Limitations
LERF provides a dense volumetric interface for zero-shot 3D language queries without region proposals or fine-tuning, with potential robotics and 3D-interaction uses.
Takeaways & Limitations
LERF requires calibrated camera matrices and NeRF-quality multi-view captures, and inherits CLIP’s bag-of-words behavior and spatial-relation difficulties.
Abstract
from arXiv · showhide
Humans describe the physical world using natural language to refer to specific 3D locations based on a vast range of properties: visual appearance, semantics, abstract associations, or actionable affordances. In this work we propose Language Embedded Radiance Fields (LERFs), a method for grounding language embeddings from off-the-shelf models like CLIP into NeRF, which enable these types of open-ended language queries in 3D. LERF learns a dense, multi-scale language field inside NeRF by volume rendering CLIP embeddings along training rays, supervising these embeddings across training views to provide multi-view consistency and smooth the underlying language field. After optimization, LERF can extract 3D relevancy maps for a broad range of language prompts interactively in real-time, which has potential use cases in robotics, understanding vision-language models, and interacting with 3D scenes. LERF enables pixel-aligned, zero-shot queries on the distilled 3D CLIP embeddings without relying on region proposals or masks, supporting long-tail open-vocabulary queries hierarchically across the volume. The project website can be found at https://lerf.io .
1. Introduction
LERF grounds CLIP language embeddings in NeRF to make 3D scenes searchable through open-ended, multi-scale natural-language queries. It produces localized, 3D-consistent relevancy maps across diverse prompts and scenes.
- NeRFs capture photorealistic 3D scenes but output a colorful density field without meaning or context.
- LERF optimizes off-the-shelf CLIP embeddings into NeRF without fine-tuning or region proposals, preserving broad semantic coverage.The supported queries include visual, abstract, textual, and long-tail concepts.
- LERF jointly optimizes a language field conditioned on position and physical scale, supervising it with multi-scale CLIP features from training-view crops.Different scales can associate one 3D location with distinct concepts such as “utensils” and “wooden spoon.”
- LERF produces more localized relevancy maps than 2D CLIP embeddings while maintaining 3D consistency across views.Queries can be made directly in the 3D field without rendering multiple views.
- LERF supports fine-grained and abstract language queries across handheld in-the-wild scenes, generating real-time, view-consistent 3D relevancy maps.The paper evaluates against LSeg and OWL-ViT and identifies potential uses in robotics, vision-language-model analysis, and 3D interaction.
2. Related Work
Related work grounds semantics in 2D images, point clouds, or NeRFs, but LERF emphasizes dense, volumetric, multi-scale language features without region proposals. This design supports hierarchical 3D text queries and preserves expressive CLIP semantics when multi-view inputs are available.
- 2. Related Work: 2D open-vocabulary methods range from zero-shot systems to models trained with segmentation data, often using pixel embeddings, decoders, or region proposals.Proposal-based methods can leverage detection data but may struggle with unlabeled hierarchical object components.
- 2. Related Work: LERF avoids region proposals by incorporating language embeddings densely in a 3D, multiscale field for hierarchical text queries.
- 2. Related Work: LERF builds a reusable 3D representation that can answer different text prompts without reconstructing the scene for each query.
- 2. Related Work: Unlike prior pixel-aligned feature distillation, LERF distills non-pixel-aligned CLIP embeddings into 3D without fine-tuning.
- 2. Related Work: Point-cloud approaches fuse CLIP crops into sparse maps, whereas LERF queries language features densely throughout the scene.
- 2. Related Work: LERF provides a dense volumetric interface that can improve the resolution and fidelity of downstream 3D-language applications when multi-view inputs are available.
3. Language Embedded Radiance Fields
LERF grounds CLIP language embeddings in a dense, multiscale 3D field within NeRF, enabling open-ended relevancy queries at arbitrary scales. Multi-view rendering and DINO regularization improve localization and smoothness, while querying supports real-time maps but depends on calibrated, high-quality captures.
- 3. Language Embedded Radiance Fields: LERF represents language as a scale-conditioned field over 3D volumes, supervised by CLIP embeddings from image crops across training views.This associates a location with distinct embeddings at different context scales, such as “utensils” versus “wooden spoon.”
- 3.1. LERF Volumetric Rendering: Language embeddings are rendered along camera rays using NeRF weights over view-dependent frustum scales, then averaged and normalized as CLIP embeddings.The frustum scale increases with focal length and sample distance, enabling pixel-aligned rendering from a volumetric field.
- 3.2. Multi-Scale Supervision: Multi-scale supervision uses precomputed CLIP image pyramids and interpolated crop embeddings to train language fields across randomly sampled ray origins and physical scales.The rendered embedding is trained to maximize cosine similarity with the interpolated ground-truth embedding.
- 3.3. DINO Regularization: DINO regularization improves relevancy-map smoothness and boundaries, especially where views are sparse or foreground and background have little geometric separation.Removing DINO produces qualitative deterioration in these regions.
- 3.6. Implementation Details: LERF’s language fields require calibrated camera matrices and NeRF-quality multiview captures, and nearby unseen surfaces can produce blurred relevancy maps.The method is implemented on Nerfacto with a multiresolution hashgrid and separate CLIP and DINO feature networks.
- 3.5. Querying LERF: LERF computes query relevancy against canonical negative phrases, selects a single scale by maximizing score, and filters samples observed by fewer than five training views.The scale heuristic evaluates 30 scales from 0 to 2 meters and assumes relevant scene parts share the same scale.
4. Experiments
LERF is evaluated on diverse in-the-wild and long-tail scenes for qualitative language querying, existence determination, and localization. Across these experiments, it supports multi-scale semantic queries and outperforms LSeg on long-tail 3D language grounding, while ablations identify benefits from DINO and multi-scale supervision.
- Experimental Setup: LERF is evaluated on 13 hand-held captured scenes spanning grocery stores, kitchens, bookstores, and posed long-tail scenes.The scenes are captured with Polycam using images at 994×738 resolution.
- Qualitative Results: LERF supports semantic queries ranging from visual properties and specific book or character names to abstract groupings such as “cartoon.”The same object can receive multiple semantic tags because the representation lacks discrete categories.
- Existence Determination: LSeg performs similarly to LERF on in-distribution labels but significantly suffers on long-tail labels in wild scenes.Figure 7 illustrates this gap: LSeg locates “glass of water” but cannot locate an out-of-distribution “egg.”
- Localization: LERF strongly outperforms LSeg in 3D localization across 72 objects in 5 scenes, while OWL-ViT outperforms LSeg but suffers relative to LERF on long-tail queries.A localization success occurs when the highest relevancy pixel lands inside the labeled bounding box.
- Ablations: Removing DINO qualitatively deteriorates relevancy-map smoothness and boundaries, especially with few surrounding views or weak foreground-background separation.The ablation presents two illustrative examples where DINO improves relevancy-map quality.
- Ablations: Single-scale CLIP supervision significantly impairs queries at different scales, including large “espresso machine” queries and “creamer pods” queries.The results imply that multi-scale training regularizes the language field across scales.
5. Limitations
LERF inherits limitations from CLIP and NeRF, including semantic false positives, weak spatial reasoning, and dependence on high-quality calibrated multi-view captures. Its single-scale query rendering can also miss context needed by some queries.
- CLIP Limitations: CLIP-like bag-of-words behavior makes “not red” similar to “red” and limits spatial reasoning between objects.LERF can also activate on visually or semantically similar distractors, such as other similarly shaped vegetables for “zucchinis.”
- NeRF and Capture Requirements: LERF requires known calibrated camera matrices and NeRF-quality multi-view captures, which are not always available or easy to obtain.Language-field quality is bottlenecked by NeRF reconstruction quality.
- NeRF and Capture Requirements: Objects near other surfaces can have blurred embeddings when side views do not expose the background without the object.This produces blurry relevancy maps similar to single-view CLIP.
- Scale Limitation: LERF renders language embeddings from a single scale for each query, although some queries such as “table” may benefit from or require multiple scales.
6. Conclusions
LERF fuses raw CLIP embeddings into NeRF as a dense, multi-scale field without region proposals or fine-tuning. It supports broad natural-language queries across diverse real-world scenes, while relevancy maps can also assign non-zero scores to regions that resemble the query.
- Conclusion: LERF fuses raw CLIP embeddings into NeRF densely and at multiple scales without region proposals or fine-tuning.The framework is described as supporting aligned multimodal encoders beyond CLIP.
- Conclusion: LERF supports broad natural-language queries across diverse real-world scenes and strongly outperforms pixel-aligned LSeg for natural-language queries.
- Conclusion: Relevancy maps are generated by thresholding similarity against canonical negative phrases, with scores below 0.5 treated as irrelevant in the project videos.
- Conclusion: Regions similar to, but not matching, a query receive non-zero relevancy scores, while longer-tail and more specific queries tend to separate more clearly from canonical phrases.The behavior can group semantically similar regions but may also produce too many relevant regions.
B. Additional qualitative results
The appendix provides additional qualitative results from scenes not shown in the main text or videos.
- Additional qualitative results: Figure 16 presents additional qualitative results from scenes not pictured in the main text or videos.
C. Numerical relevancy scores
LERF relevancy scores vary with query specificity: descriptive visual-semantic prompts score higher, while abstract or poorly observed queries score lower. CLIP ambiguity can also produce incorrect activations.
- Visual similarity can divert relevancy from the queried object to unrelated but similar regions.Examples include bell peppers activating on jalepenos and portafilter activating on a grinder spout.
D. Convergence speed
LERF produces usable relevancy maps early in optimization, while continued training refines fine-grained and small-object queries. Convergence is faster for common semantics and regions with broader multi-view coverage.
- Usable 3D language embeddings emerge within the first thousand optimization steps, although fine-grained queries and small objects improve with further training.Relevancy maps are compared at 1k, 2k, 6k, and 30k steps.
- Common semantics and expansive multi-view regions converge faster than fine-grained properties.“Blue dish soap” converges faster than “bath toys.”
E. Experiment details
The experiments use labeled views and long-tail scene labels, while additional analyses examine prompt wording, CLIP’s bag-of-words behavior, and geometry-related degradation.
- More specific prompt wording can improve relevancy-map quality for some objects.
- The existence experiment uses custom long-tail labels, with positive labels assigned per scene.
- CLIP may behave as a bag of words, causing adjectives to be incorporated improperly into queries.
- Poor NeRF geometry, including floaters and incomplete geometry, can make rendered CLIP embeddings unreliable.
- Limited geometric separation can blur query activations between objects and foreground-background regions.The toaster example is fuzzier because few viewing angles were captured.
F. Detailed Illustrations of Limitations
LERF’s limitations arise from inherited CLIP ambiguities and NeRF geometry failures, while prompt specificity can improve localization. Additional figures and label tables document these behaviors and experimental labels.
- LERF inherits CLIP limitations including language ambiguity and prompt sensitivity, as well as NeRF geometry limitations.
- Visually similar regions can produce false-positive relevancy activations for queries such as portafilter and refrigerator.
- CLIP’s bag-of-words behavior can make added adjectives shift activation toward incorrect regions.Examples include “mug handle” versus “handle” and “coffee spill” versus “spill.”
- Unreliable NeRF geometry degrades language localization for reflective and transparent objects.Holes can cause multi-view embeddings to average incorrectly, while transparent objects may receive weight from the opaque background.
- More descriptive prompts can refine or shift relevancy toward the intended object.“Blue dish soap” shifts activation from a pump soap bottle to the correct object compared with “Dish soap.”
- Insufficient geometric separation can blur relevancy maps into surrounding objects.The effect occurs when foreground objects appear in front of the background across most views.
- The supplementary material includes additional scenes and tables listing labels used in detection and existence experiments.