A Comparative Analysis of Four Semantic Segmentation Models for Classification of Building Structural Elements in Highly Occluded Environments
Keywords: BIM, Scan-to-BIM, Semantic Segmentation, Structural Model, Point-Based Networks, Point Cloud
Abstract. Semantic segmentation of point clouds plays a critical role in automatic Scan-to-BIM workflows. This study evaluates the performance of two Multi-Layer-Perceptron (MLP)-based deep learning networks (PointNet++, PointNeXt-XL) and two transformer-based networks (Point Transformer V1 (PTv 1), Point Transformer V3 (PTv 3)) for the identification of building structural elements in highly occluded environments. The models are trained and tested on three LiDAR datasets of reinforced concrete structures including office buildings and a multi-storey carpark. Results show that transformer-based networks significantly outperform MLP-based architectures under heavy occlusions. Within the transformer-based models, PTv 1 achieved the highest Overall Accuracy (OA) at 92.62% while PTv 3 provided the best balance across all classes, particularly for beam and clutter classes due to its larger receptive field and enhanced geometric encoding. Because segmentation approaches can compensate for errors in ceiling, floor, and column class identification, we recommend PTv 3 for Scan-to-BIM applications focused on structural modelling.
