Enhancing semantic segmentation of construction scenes using IFC-derived synthetic point clouds for geometric Digital Twins
Keywords: Construction, Digital Twin, Virtual Laser Scanning, Semantic Segmentation, BIM, point clouds
Abstract. Accurate semantic segmentation of as-built point clouds is foundational for Digital Twin Construction, yet training deep learning models is constrained by data scarcity in the AEC sector. This study presents a method to create semantically labeled synthetic point clouds from IFC models, and employs geometric augmentation to bridge the domain shift to real-world LiDAR data. We developed an automated pipeline to generate semantically labeled synthetic point clouds directly from IFC models, and conducted a systematic hyperparameter search, testing the impact of point-level Gaussian noise (Jitter Standard Deviation) and maximum displacement (Clipping Range) on the Oneformer3D model’s ability to segment complex Mechanical, Electrical, and Plumbing (MEP) components. Our results demonstrate that this augmentation approach significantly enhances segmentation performance compared to baseline model, particularly in real-world environments. The optimal configuration achieved a synthetic mIoU of 67.31% (up from 62.52%) and a real-world mIoU of 32.75% (improving significantly upon the 14.66% baseline). These results confirm that balancing high noise variance (for domain generalization) with strict displacement control (to retain feature integrity) is key to successful synthetic-to-real domain transfer. This approach offers a scalable alternative for generating the necessary training data for automated geometric Digital Twin realization.
