Memory-Efficient 3D Scene Understanding via Quantized Multi-View Features
Keywords: 3D scene understanding, foundation models, feature quantization, large-scale point clouds, 2D-to-3D distillation
Abstract. Projecting high-dimensional representations derived with 2D foundation models onto 3D point clouds via multi-view aggregation has emerged as a powerful paradigm for 3D scene understanding. However, the conventional strategy of storing dense feature vectors - derived by 2D foundation models at each point - introduces a substantial memory overhead that scales linearly with the scene size, thereby limiting scalability and hindering practical deployment in large-scale settings. In this work a memory-efficient framework is proposed to address this limitation through feature quantization. Specifically, projected 2D features are clustered into a compact set of prototypical embeddings, enabling each 3D point to be represented by a single discrete index instead of a full high-dimensional descriptor. This representation drastically reduces memory requirements by orders of magnitude while preserving the semantic richness of the original features. The proposed quantization framework is validated on diverse point clouds, demonstrating that the representations retain strong performance in downstream tasks while significantly improving computational efficiency.
