<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">ISPRS-Annals</journal-id>
<journal-title-group>
<journal-title>ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences</journal-title>
<abbrev-journal-title abbrev-type="publisher">ISPRS-Annals</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2194-9050</issn>
<publisher><publisher-name>Copernicus Publications</publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/isprs-annals-XII-4-W1-2026-227-2026</article-id>
<title-group>
<article-title>Closing the Loop: Perceptually-Guided Iterative Generation of LoD2.0 3D Building Meshes for Building Digital Cousin development</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Liao</surname>
<given-names>Lingfeng</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhao</surname>
<given-names>Chenbo</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Sekimoto</surname>
<given-names>Yoshihide</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Ogawa</surname>
<given-names>Yoshiki</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>Center for Spatial Information Science, The University of Tokyo, Chiba, Japan</addr-line>
</aff>
<aff id="aff2">
<label>2</label>
<addr-line>Graduate School of Social Data Science, Hitotsubashi University, Tokyo, Japan</addr-line>
</aff>
<pub-date pub-type="epub">
<day>28</day>
<month>09</month>
<year>2026</year>
</pub-date>
<volume>XII-4/W1-2026</volume>
<fpage>227</fpage>
<lpage>234</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Lingfeng Liao et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/227/2026/isprs-annals-XII-4-W1-2026-227-2026.html">This article is available from https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/227/2026/isprs-annals-XII-4-W1-2026-227-2026.html</self-uri>
<self-uri xlink:href="https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/227/2026/isprs-annals-XII-4-W1-2026-227-2026.pdf">The full text article is available as a PDF file from https://isprs-annals.copernicus.org/articles/XII-4-W1-2026/227/2026/isprs-annals-XII-4-W1-2026-227-2026.pdf</self-uri>
<abstract>
<p>This study proposes an advanced generative framework for creating digital building representations, building upon our previously developed autoregressive Building Digital Cousin (BDC) generator. Previous approaches to building reconstruction have relied heavily on extensive visual data references, which often lead to significant data deficiencies. Our earlier method partially addressed these limitations by incorporating generative modeling based on textual and image targets. However, this approach still suffered from instability during multi-round generation and from limitations associated with manually configured conditioning. This proposal advances the configuration by introducing a simple but novel perceptual evaluation index, Building Mesh Quality Index (BMQI), for building instances. Additionally, an expanded vocabulary for vertex discretization and a latent image-based conditioning mechanism are introduced to refine the previous pipeline. By setting the generation target to LoD2.0 without minor-level details, experiments on the PLATEAU dataset demonstrate the capability of our model to generate a wide range of building models that conform to the image description while maintaining outstanding perceptual quality in more than 99% of cases. An average improvement of 13% in geometric proximity and a BMQI close to 0.99 further demonstrate the effectiveness of our method in addressing previous deficiencies related to data dependency and appearance conformity in urban building reconstruction.</p>
</abstract>
<counts><page-count count="8"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>