Source-linked AI summary
Linking Points With Labels in 3D: A Review of Point Cloud Semantic Segmentation
Yuxing Xie, Jiaojiao Tian, Xiao Xiang Zhu
TL;DR
PCSS research spans diverse applications and must handle complex 3D point-cloud data, while existing surveys lacked sufficient detail amid rapid deep-learning development. This paper reviews acquisition, benchmarks, PCS and PCSS methods, and open issues, finding strong benchmark performance from deep-learning methods but limited cross-dataset comparability and no standard public neural network.
Problem
PCSS supports remote sensing, computer vision, robotics, and object recognition, but the field needed an updated synthesis as earlier reviews lacked detail and deep learning rapidly expanded.
Method
The paper reviews point-cloud acquisition and evolution, benchmarks, traditional and advanced PCS/PCSS algorithms, and issues across the field.
Results
Deep-learning methods rank highly on most benchmark evaluations, but methods have been tested on limited and dissimilar datasets and no standard neural network is publicly available.
Takeaways & Limitations
Selecting an optimal practical PCSS approach remains difficult because existing methods are evaluated on limited, dissimilar datasets.
Takeaways & Limitations
Public benchmark datasets remain insufficient for PCSS tasks, despite the availability of large datasets containing hundreds of millions of points.
Abstract
from arXiv · showhide
3D Point Cloud Semantic Segmentation (PCSS) is attracting increasing interest, due to its applicability in remote sensing, computer vision and robotics, and due to the new possibilities offered by deep learning techniques. In order to provide a needed up-to-date review of recent developments in PCSS, this article summarizes existing studies on this topic. Firstly, we outline the acquisition and evolution of the 3D point cloud from the perspective of remote sensing and computer vision, as well as the published benchmarks for PCSS studies. Then, traditional and advanced techniques used for Point Cloud Segmentation (PCS) and PCSS are reviewed and compared. Finally, important issues and open questions in PCSS studies are discussed.
I. MOTIVATION
PCSS extends semantic segmentation to 3D point clouds, which are increasingly accessible through diverse acquisition technologies and support applications requiring object- or class-level information. This review updates earlier surveys by organizing point-cloud acquisition, benchmarks, PCS and PCSS algorithms, and open issues amid rapid deep-learning progress.
- PCSS assigns semantic labels to regularly or irregularly distributed 3D points rather than 2D pixels.
- PCS groups points by geometric or spectral similarity without semantic information and can serve as a presegmentation step in PCSS.
- Object- and class-level point-cloud information supports urban planning, forest monitoring, robotics mapping, autonomous driving, and HD-map construction.
- The review addresses limited detail in earlier surveys by covering point-cloud acquisition and benchmarks, traditional and advanced PCS/PCSS algorithms, and current issues.
- Point clouds arise from image-derived methods, LiDAR, RGB-D cameras, and SAR systems, whose data features and application ranges differ.
2) LiDAR point cloud:
LiDAR measures target distance using laser energy and supports point clouds spanning airborne, terrestrial, mobile, and unmanned platforms. These variants trade off density, coverage, flexibility, spectral information, and acquisition context across applications.
- LiDAR emits laser pulses and measures their travel time to determine distances to surveyed objects.
- ALS provides relatively low-density airborne point clouds, while TLS produces very dense, high-quality 3D models for small urban, forest, heritage, and artwork sites.
- MLS is commonly vehicle-mounted and is strongly associated with autonomous-driving research and HD-map generation.
- ULS uses drones or unmanned vehicles to collect denser, more accurate data than ALS with greater operational flexibility.
- LiDAR surveys require GNSS and IMU data to match moving-platform measurements, and LiDAR is also used as ground truth for evaluating other point clouds.
4) SAR point cloud:
SAR-derived point clouds offer complementary geometric and temporal information, especially for urban structures, but remain less widely used and face accuracy and outlier limitations. Point-cloud characteristics and applications have evolved with denser, larger-volume, and multi-source data.
- InSAR techniques such as TomoSAR and PSI generate point clouds from SAR imagery for surface-deformation or elevation-related remote sensing.
- TomoSAR positioning accuracy is typically about 1 m, compared with about 0.1 m for ALS LiDAR, although correction can reach decimeter-level accuracy.
- TomoSAR provides rich façade information and can support detailed building reconstruction, whereas incoherent objects such as trees may not be reconstructed.
- SAR is currently the only spaceborne sensor described as providing fourth-dimensional temporal deformation information alongside geometric and material façade properties.
- InSAR accuracy is affected by anisotropic elevation error and ghost scatterers produced by multiple scattering, while global SAR data support future expansion.
- Point clouds have progressed from sparse data that could not represent land cover at object level toward dense, large-volume, and multi-source data requiring practicable algorithms.
C. Point cloud application
PCS and PCSS studies select data and algorithms according to application environments, with LiDAR dominating current work and benchmarks enabling evaluation. Dataset coverage remains uneven, especially for image-derived and early benchmark data.
- Reviewed studies are classified by point-cloud type and working environment, including urban, forest, industry, and indoor settings.
- LiDAR is the most commonly used PCS data source, particularly for buildings in urban environments and trees in forests.
- Image-derived point clouds appear frequently in real-world scenarios, but limited annotated benchmarks constrain PCS and PCSS research on them.
- RGB-D data are mainly used indoors because of close sensing range; plane segmentation dominates PCS, while several RGB-D benchmarks support deep-learning PCSS evaluation.
- InSAR point clouds have relatively few studies but show potential for urban monitoring, especially building-structure segmentation.
- Benchmark datasets support algorithm development, evaluation, and comparison, but early datasets had shortcomings and mainstream datasets remain concentrated in LiDAR or RGB-D sensing.
- Semantic3D.net is a representative large-scale outdoor TLS dataset containing over four billion labeled 3D points across eight urban-object classes.
2) Stanford Large-scale 3D Indoor Spaces Dataset (S3DIS):
S3DIS is a large-scale indoor RGB-D benchmark with over 215 million instance-level annotated points across six regions in three buildings. The surrounding overview also organizes point-cloud applications and segmentation studies by acquisition type and environment.
- S3DIS: S3DIS contains over 215 million points across more than 6,000 m2 in six indoor regions from three buildings.Its main areas are educational and office spaces.
- S3DIS: Its annotations are instance-level, covering structural and movable elements divided into 13 classes.
- Point-cloud overview: Table I compares point-cloud types by density, advantages, disadvantages, and applications.
- PCS and PCSS applications: Table II classifies PCS and PCSS studies by acquisition type and working environment, including urban, forest, industry, and indoor settings.Methods are represented using abbreviations, including region growing, Hough Transform, RANSAC, clustering, machine learning, and deep learning.
- PCS and PCSS applications: The reviewed studies span building planes, urban scenes, tree and forest structures, planes, buildings, trees, and related PCSS tasks.The listed methods range from region growing, Hough Transform, clustering, and RANSAC to machine learning and deep learning.
InSAR
The supplied passages identify a sequence of studies associated with the section, but provide no substantive description of the InSAR dataset or its properties.
- InSAR: The passage lists three associated studies as [129] (2005/HT), (2015/R), and (2018/R).
- InSAR: The surrounding dataset discussion describes an influential ALS benchmark whose labeled point cloud is divided into nine classes for algorithm evaluation.
4) Paris-Lille-3D:
Paris-Lille-3D is a recent MLS benchmark for PCSS, while ScanNet provides a complementary voxel-based indoor RGB-D benchmark. The reviewed PCS literature also distinguishes segmentation from semantic labeling.
- Paris-Lille-3D: Paris-Lille-3D contains more than 140 million labeled points across 50 urban object classes along 2 km of streets in Paris and Lille.It can also support autonomous-vehicle research.
- Paris-Lille-3D: Because Paris-Lille-3D was recent, only a few validated results were available on its related website.
- ScanNet: ScanNet v2 contains 1,513 annotated scans with approximately 90% surface coverage and 20 classes of annotated 3D voxelized objects.Unlike the other benchmarks, it labels voxels rather than points or objects.
- PCS background: Traditional PCS groups raw 3D points into non-overlapping regions using hand-crafted geometric and statistical rules, without supervised prior knowledge.The resulting regions have no strong semantic information.
A. Edge-based
The review covers edge-based, region-growing, and model-fitting approaches for point-cloud segmentation, emphasizing their operating principles and practical limitations. Edge methods are fast but scene-sensitive, region growing depends on tuned criteria and seeds, and Hough-based fitting faces computational costs despite robust 3D variants.
- A. Edge-based: Edge-based PCS detects boundaries from rapid intensity changes and groups points inside those boundaries into final segments.Its standard pipeline has edge detection followed by point grouping.
- A. Edge-based: Edge-based methods are fast and simple but perform reliably mainly on low-noise, evenly dense scenes and are rarely used for dense or large-area datasets.Some variants are limited to range images, and disconnected 3D edges may not directly define closed segments.
- B. Region growing: Region growing merges spatially close points or regions with similar surface properties, using seed selection, growth units, and similarity criteria.Growth units may be single points, voxel or octree regions, or hybrid units.
- B. Region growing: Region-growing accuracy depends on predefined growth criteria and seed locations, which must be adjusted across datasets.These methods are also computationally intensive and may require data reduction to balance accuracy and efficiency.
- C. Model fitting: Model fitting segments parameterized geometric shapes by matching point clouds to primitives such as planes, spheres, and cylinders, commonly using Hough Transform or RANSAC.
- C. Model fitting: 3D KHT performed better than previous Hough Transform techniques, including RHT, for plane detection and was robust to noise and irregularly distributed samples.
2) RANSAC:
RANSAC fits predefined geometric models by generating hypotheses from random samples and selecting the best-supported model. It is efficient and robust to noise, but its nondeterminism can produce spurious surfaces.
- RANSAC procedure: RANSAC has two phases: generate model hypotheses from random samples, then evaluate them against the data.Models must be manually defined or selected before hypothesis generation.
- RANSAC procedure: For plane segmentation, RANSAC estimates a plane from three non-collinear points and represents it with parameters [a, b, c, d]T.Three non-collinear points determine a plane.
- Hypothesis evaluation: RANSAC selects the most probable hypothesis by minimizing a loss over the data, using an error function such as geometric distance.The selection is formulated as an optimization problem.
- Advantages: RANSAC avoids complex optimization and high memory requirements while processing data with substantial noise and outliers.Compared with Hough Transform methods, efficiency and the percentage of successfully detected objects are identified as major advantages.
- Applications: RANSAC supports segmentation of planes and more complex primitives, including spheres, cylinders, cones, and tori.Applications include building façades, building roofs, indoor scenes, and cylinder objects.
- Limitations and improvements: Because RANSAC is nondeterministic, it can detect spurious surfaces representing models that do not exist in reality.Soft-threshold voting and NDT-cell-based improvements were proposed to reduce this problem.
3) Mean-shift:
Mean-shift is a nonparametric clustering method that avoids predefining the cluster count, unlike K-means. Its tendency toward oversegmentation often makes it useful as a presegmentation step before later partitioning or refinement.
- Limitation: Because both cluster number and cluster shape are unknown, mean-shift produces highly probable oversegmented results.This behavior motivates its use before partitioning or refinement.
- Graph-based PCS: Graph-based PCS methods construct neighborhood graphs and partition them, including min-cut formulations for outdoor urban objects and ALS roof segmentation.Points serve as graph nodes connected by neighborhood edges, such as k-nearest-neighbor or 3D Voronoi relationships.
- Contextual PCSS: CRF and supervised MRF methods are used as contextual models or postprocessing stages in PCSS.They address contextual information that individual point classifiers do not capture.
- Oversegmentation and presegmentation: Presegmentation into voxels or supervoxels reduces point-cloud data volume before computationally expensive processing, but fixed-resolution VCCS may yield poor boundaries in non-uniform density.VCCS voxelizes the cloud with an octree and applies K-means clustering for supervoxel segmentation.
B. Deep learning
Deep learning-based PCSS methods address unordered and irregular point clouds through multiview, voxel, or point-based representations. Direct point processing avoids some structural losses, while deep learning has improved accuracy but remains limited by interpretability and data requirements.
- Deep learning: Deep learning-based PCSS approaches transform point clouds into multiview, voxel, or point-based representations because standard convolutions target ordered raster data.The three categories depend on the format ingested by neural networks.
- 1) Multiview-based: Multiview methods project 3D data into 2D images for image-based learning, then restore semantic results to the 3D point cloud.SnapNet preprocesses points, generates mesh-based RGB and depth images from virtual cameras, segments them, and projects labels back.
- 1) Multiview-based: Multiview methods lose geometric structure through 2D approximation and require viewpoints covering all points, making large complex scenes difficult to process.These limitations have restricted the use of multiview architectures for PCSS.
- 2) Voxel-based: Voxelization enables 3D convolutions by converting unordered points into a structured grid, but it lowers resolution and stores free or unknown spaces.The resulting representation can require substantial computation and memory.
- 2) Voxel-based: SegCloud combines voxelization, a 3D fully convolutional network, trilinear interpolation, and fully connected CRF regularization to produce point labels.The pipeline predicts downsampled voxel labels, transfers them to points, and regularizes the results.
- 3) Directly process point cloud data: Direct point-based methods bind point-cloud canonicalization to the network architecture, avoiding separate multiview or voxel preprocessing.PointNet uses per-point MLPs, symmetric max pooling for global features, and concatenated local and global features for segmentation.
- 3) Directly process point cloud data: PointNet remains a PCSS baseline, while later models add local neighborhoods, long-range context, convolutions, or multimodal fusion.Examples include PointNet++, 3P-RNN, annular convolution, PointCNN, and frameworks combining images with 3D points.
- 5) Regular machine learning vs. deep learning: Deep learning has boosted 3D PCSS accuracy and handles large datasets efficiently without handcrafted feature design, but its internal decisions remain poorly interpretable and data-limited.These limitations matter especially in applications requiring high safety or stability.
C. Hybrid methods
Hybrid PCSS methods first create segments or superpoints and then classify those units, using presegmentation to reduce data volume and provide local features. Feature design remains a central accuracy–efficiency trade-off, while contextual and graph-based models still require further improvement or evaluation.
- C. Hybrid methods: Hybrid segment-wise PCSS methods oversegment or partition point clouds before applying semantic segmentation to segments rather than individual points.Presegmentation generally reduces data volume and supplies local features.
- C. Hybrid methods: Supervoxels, region growing, Hough-transform patches, and superpoint graphs are examples of presegmentation used in hybrid PCSS pipelines.These approaches combine nonsemantic PCS or oversegmentation with supervised or deep learning-based semantic prediction.
- Features: Feature selection trades algorithm accuracy against efficiency, with PCSS features differing by neighborhood selection, extraction scale, and application.The review identifies feature design, selection, and application as major differences among PCS and PCSS methods.
- Features: Local features remain a significant improvement target after PointNet because its original formulation does not represent neighboring-point structure.Subsequent methods address this through hierarchical, pooling, convolutional, or recurrent mechanisms.
- Future directions: Presegmentation can provide local features naturally, and the review identifies nonsemantic presegmentation as a possible future direction for PCSS.This direction is presented conditionally on neural-network architectures becoming more stable.
- Contextual information: Contextual models are widely used for supervised PCSS, often as smoothing postprocessing, but contextual segmentation still has room for improvement.Several deep learning methods also employ contextual segmentation.
- Contextual information: Graph neural networks have shown excellent PCSS performance in reported studies, but more research is required to evaluate their performance.The review presents GNNs as a promising but still developing direction.
B. Remote sensing meets computer vision
Remote sensing and computer vision share point-cloud research but differ in data, evaluation priorities, and application demands. The review highlights limited benchmark coverage, unavoidable noise, and the need for more robust, broadly applicable PCSS methods.
- Computer vision develops accuracy-oriented algorithms, whereas remote sensing applies methods across datasets but often cannot adopt them directly.
- Remote sensing prioritizes object-specific accuracy, while generic computer vision often emphasizes overall accuracy.Building accuracy is especially important for urban monitoring.
- Remote sensing requires large-area data with complex, specific categories that exceed the scope of many small-area computer vision benchmarks.Agricultural applications may require separating vegetation into species.
- Noise and outliers remain unavoidable in remote sensing, while current methods lack noise adaptation and robustness.The review identifies denoising and sensor-theory integration as future research directions.
- Benchmark datasets have grown substantially, yet they remain insufficient for PCSS because real-world object categories and environments are more varied.Semantic3D.net, for example, covers only one kind of city.
- Public benchmarks underrepresent image-derived and remote-sensing point clouds, including airborne 2.5D, satellite photogrammetric, InSAR, and multi-source data.Only the Vaihingen dataset is identified as a published benchmark for remote-sensing PCSS tasks.
- Deep learning methods rank highly in benchmark evaluations, but limited and dissimilar datasets make practical method selection difficult.The review notes that no standard neural network for PCSS is publicly available.