<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Annals</journal-id>
<journal-title-group>
<journal-title>ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Annals</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9050</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-annals-XII-4-W1-2026-9-2026</article-id>
<title-group>
<article-title>Memory-Efficient 3D Scene Understanding via Quantized Multi-View Features</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Alami</surname>
<given-names>Ashkan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Remondino</surname>
<given-names>Fabio</given-names>
<ext-link>https://orcid.org/0000-0001-6097-5342</ext-link>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>3D Optical Metrology (3DOM) unit, Bruno Kessler Foundation (FBK), Trento, Italy</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>University of Trento, Department of Information Engineering and Computer Science (DISI), Trento, Italy</addr-line>
</aff>
<pub-date pub-type="epub">
<day>28</day>
<month>09</month>
<year>2026</year>
</pub-date>
<volume>XII-4/W1-2026</volume>
<fpage>9</fpage>
<lpage>16</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Ashkan Alami</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/9/2026/isprs-annals-XII-4-W1-2026-9-2026.html">This article is available from https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/9/2026/isprs-annals-XII-4-W1-2026-9-2026.html</self-uri>
<self-uri xlink:href="https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/9/2026/isprs-annals-XII-4-W1-2026-9-2026.pdf">The full text article is available as a PDF file from https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/9/2026/isprs-annals-XII-4-W1-2026-9-2026.pdf</self-uri>
<abstract>
<p>Projecting high-dimensional representations derived with 2D foundation models onto 3D point clouds via multi-view aggregation has emerged as a powerful paradigm for 3D scene understanding. However, the conventional strategy of storing dense feature vectors - derived by 2D foundation models at each point - introduces a substantial memory overhead that scales linearly with the scene size, thereby limiting scalability and hindering practical deployment in large-scale settings. In this work a memory-efficient framework is proposed to address this limitation through feature quantization. Specifically, projected 2D features are clustered into a compact set of prototypical embeddings, enabling each 3D point to be represented by a single discrete index instead of a full high-dimensional descriptor. This representation drastically reduces memory requirements by orders of magnitude while preserving the semantic richness of the original features. The proposed quantization framework is validated on diverse point clouds, demonstrating that the representations retain strong performance in downstream tasks while significantly improving computational efficiency.</p>
</abstract>
<counts><page-count count="8"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>