Next Article in Journal
Information Sustainability Beyond Digital Access: Machine Learning Evidence from Local Media Ecosystems in Ecuador
Previous Article in Journal
Entrepreneurship Education and Entrepreneurial Intention Among University Students: The Mediating Roles of Entrepreneurial Self-Efficacy and Motivation
Previous Article in Special Issue
A Stochastic Multi-Objective Model for Optimal Design of Electronic Waste Reverse Supply Chain
 
 
Article
Peer-Review Record

Context-Responsive Building Footprint Generation via Conditional Inpainting Using Latent Diffusion Models

Sustainability 2026, 18(8), 3987; https://doi.org/10.3390/su18083987
by Eunseok Jang and Kyunghwan Kim *
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Reviewer 3:
Sustainability 2026, 18(8), 3987; https://doi.org/10.3390/su18083987
Submission received: 11 March 2026 / Revised: 6 April 2026 / Accepted: 13 April 2026 / Published: 17 April 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

While numerous recent studies have explored the applications of Latent Diffusion Models (LDM)/Stable Diffusion in architectural form generation and landscape design, this manuscript only provides a detailed analysis of research on models such as GANs and VAEs. The review of state-of-the-art applications of diffusion models in architecture remains insufficient, and the differentiated contributions of this study from existing diffusion model-based architectural research are not clearly articulated. It is recommended to supplement and comparatively analyze relevant research findings from the past three years.

The manuscript mentions that the context-responsive design of building footprints should align with urban fabric, yet the core evaluation dimensions of "context responsiveness" in this study are not explicitly defined and are only scattered in the qualitative analysis. It is suggested to clarify the quantitative and qualitative evaluation criteria for context responsiveness in the Introduction or Theoretical Background section to enhance the theoretical rigor of the research.

The study selected sites in Korean new towns and areas surrounding landmark buildings as training data, which feature relatively regular urban fabric but lack coverage of irregular sites and non-grid fabric sites in old urban areas. It is recommended to elaborate on the limitations of the dataset and supplement test results from a small number of irregular sites to verify the model’s generalization ability. Additionally, the screening criteria for the "expert-designed cases" in the training data are not specified; the basis for case selection should be supplemented.

The study set the guidance scale (s) to 1.2, dropout rate to 10%, batch size to 16, and other hyperparameters, only stating that these "optimal values were determined through repeated experiments" without presenting the hyperparameter tuning process and comparative results. It is recommended to supplement the results of hyperparameter sensitivity analysis to improve the reproducibility of the experimental design.

The qualitative analysis of the three selected test sites only focuses on the morphological alignment with site context, without comparing the generated footprints with existing/optimal design schemes of the actual sites, nor analyzing the feasibility of the model’s generated results for subsequent architectural design. It is suggested to supplement a comparative analysis with practical design schemes to enhance the persuasiveness of the qualitative results.

Figures 5 and 6 present comparisons of generated results and heatmaps, yet the manuscript only provides figure titles and brief descriptions without a detailed explanation of the meaning of colors, thresholds, and annotations in the figures. It is recommended to improve the figure legends and annotations to enhance readability.

The citation formats of some references are inconsistent and need to be standardized in accordance with the journal’s requirements.

Formatting errors exist in some mathematical formulas and require careful proofreading and correction.

Author Response

Dear Reviewer,

We would like to thank you for your careful review and thoughtful comments on this manuscript. Please note that changes made to reflect the opinions of you and other reviewers are highlighted in red font in the revised manuscript. The changes made to reflect your opinion are as follows.

  1. While numerous recent studies have explored the applications of Latent Diffusion Models (LDM)/Stable Diffusion in architectural form generation and landscape design, this manuscript only provides a detailed analysis of research on models such as GANs and VAEs. The review of state-of-the-art applications of diffusion models in architecture remains insufficient, and the differentiated contributions of this study from existing diffusion model-based architectural research are not clearly articulated. It is recommended to supplement and comparatively analyze relevant research findings from the past three years.

> In the revised manuscript, Section 2.3 (Deep Learning Approaches) has been expanded to incorporate recent studies on diffusion models published between 2023 and 2024, including applications in 3D urban block generation (Zhuang et al., 2024) and multi-conditional floor plan generation (Zeng et al., 2024; Shabani et al., 2023). In addition, the manuscript now more clearly articulates the differentiated contribution of this study. Specifically, we emphasize that the proposed approach applies a Latent Diffusion Model (LDM) to 2D building footprint generation conditioned on complex geometric urban fabrics, such as surrounding street networks and adjacent buildings, which remain less explored in existing diffusion-based architectural research.

  1. The manuscript mentions that the context-responsive design of building footprints should align with urban fabric, yet the core evaluation dimensions of "context responsiveness" in this study are not explicitly defined and are only scattered in the qualitative analysis. It is suggested to clarify the quantitative and qualitative evaluation criteria for context responsiveness in the Introduction or Theoretical Background section to enhance the theoretical rigor of the research.

> To improve the theoretical rigor of the study, three qualitative evaluation criteria grounded in architectural and urban morphology theories have been introduced in Section 5.2 (Qualitative Analysis) of the revised manuscript. Based on these criteria, a structured qualitative analysis was conducted. The proposed criteria include: (1) geometric alignment with the public street edge, (2) morphological modulation with adjacent buildings, and (3) spatial differentiation of public and private domains.

  1. The study selected sites in Korean new towns and areas surrounding landmark buildings as training data, which feature relatively regular urban fabric but lack coverage of irregular sites and non-grid fabric sites in old urban areas. It is recommended to elaborate on the limitations of the dataset and supplement test results from a small number of irregular sites to verify the model’s generalization ability. Additionally, the screening criteria for the "expert-designed cases" in the training data are not specified; the basis for case selection should be supplemented.

> In Section 4.1.1 (Data Preparation), we have clarified the limitations of the dataset and explicitly stated that new-town districts were intentionally selected to provide clearly structured geometric order for the initial model training. We have also specified the selection criteria for the “expert-designed cases,” defining them as relatively recent buildings that comply with contemporary regulatory constraints and align with surrounding street axes. In addition, the Conclusion has been revised to acknowledge the limitation of the current dataset and to propose future work that includes expanding the dataset to irregular urban fabrics in order to further evaluate the model’s generalization capability.

  1. The study set the guidance scale (s) to 1.2, dropout rate to 10%, batch size to 16, and other hyperparameters, only stating that these "optimal values were determined through repeated experiments" without presenting the hyperparameter tuning process and comparative results. It is recommended to supplement the results of hyperparameter sensitivity analysis to improve the reproducibility of the experimental design.

> In Sections 4.2.3 and 4.4.1, we have clarified the rationale for the selected hyperparameters. The dropout rate of 10% is now described as being motivated by the CFG study of Ho and Salimans (2022), and a sensitivity analysis of the guidance scale (s) has been added to explain its effect on geometric distortion and spatial alignment. Additional justification for other hyperparameters, including batch size and learning rate, has also been provided based on hardware constraints and commonly adopted practices.

  1. The qualitative analysis of the three selected test sites only focuses on the morphological alignment with site context, without comparing the generated footprints with existing/optimal design schemes of the actual sites, nor analyzing the feasibility of the model’s generated results for subsequent architectural design. It is suggested to supplement a comparative analysis with practical design schemes to enhance the persuasiveness of the qualitative results.

> The generated building footprints are presented alongside the ground-truth (GT) designs in Figure 6 to enable direct comparison. In addition, the qualitative analysis in Section 5.2 has been refined to further examine how the generated results correspond to architectural features observed in real-world design schemes, such as setback distances and mass orientation.

  1. Figures 5 and 6 present comparisons of generated results and heatmaps, yet the manuscript only provides figure titles and brief descriptions without a detailed explanation of the meaning of colors, thresholds, and annotations in the figures. It is recommended to improve the figure legends and annotations to enhance readability.

> We have improved the readability of the figures and added a detailed description of the heatmap construction process in Section 5.1.1, along with a clear explanation of the meaning of the color gradient.

  1. The citation formats of some references are inconsistent and need to be standardized in accordance with the journal’s requirements. Formatting errors exist in some mathematical formulas and require careful proofreading and correction.

> To ensure consistency with both the actual implementation of the deep learning architecture and the original theoretical formulations, all equations have been carefully reviewed and revised. The main corrections are summarized as follows:

Equation (1) (U-Net input tensor for conditional inpainting): The order of VAE encoding and masking has been corrected to reflect the process described in Section 4.2.1, where masking with neutral padding is applied in pixel space prior to encoding. The notation of the mask and input tensor has also been clarified to align with the latent-space resolution.

Equation (2) (Classifier-Free Guidance, CFG): A parameter notation (θ) has been introduced to distinguish the learned noise prediction network from the actual Gaussian noise, and the final noise prediction formulation has been revised to be consistent with the standard notation in Ho and Salimans (2022).

Equation (3) (Pixel-wise mean information entropy): The use of the summation index has been corrected, and the probability variable notation has been revised to ensure consistency with the formal definition of Shannon entropy.

Regarding reference formatting, some inconsistencies remain due to the recent revisions. These will be carefully standardized in accordance with the journal’s requirements with the assistance of the journal’s editorial staff.

Thank you again for your time and valuable feedback.

Reviewer 2 Report

Comments and Suggestions for Authors

The manuscript addresses an interesting topic, but the empirical validation and methodological rigor are insufficient to support publication. I recommend rejection.

  1. The study is based on a limited and representative dataset, both in size and diversity. The reliance on rotational augmentation does not compensate for this limitation, as it does not introduce new urban conditions. Consequently, the generalizability of the findings remains uncertain.
  2. The training strategy raises concerns about redundancy and potential overfitting. The rapid convergence and lack of validation on more heterogeneous and unseen contexts suggest that the model may be learning dataset-specific regularities rather than transferable context-response principles.
  3. The experimental design is too narrow, relying on a single ablation between a four-channel model and a boundary-only baseline. The absence of intermediate configurations and comparisons with alternative methods prevents an assessment of the actual contribution of the proposed approach.
  4. The evaluation framework is inadequate, as it relies on entropy. While useful for measuring variability, entropy does not capture architectural quality and urban performance, yet the manuscript draws conclusions that go beyond what this metric can support.
  5. The results lack statistical and qualitative rigor. No statistical significance analysis is provided and the qualitative assessment is based on selected examples without a systematic and objective evaluation protocol, introducing potential bias.
  6. Methodological choices are not justified, reproducibility is incomplete and the manuscript overextends its claims, particularly regarding sustainability and urban performance, which are not evaluated.

Author Response

Dear Reviewer,

We would like to thank you for your careful review and thoughtful comments on this manuscript. Please note that changes made to reflect the opinions of you and other reviewers are highlighted in red font in the revised manuscript. The changes made to reflect your opinion are as follows.

  1. The study is based on a limited and representative dataset, both in size and diversity. The reliance on rotational augmentation does not compensate for this limitation, as it does not introduce new urban conditions. Consequently, the generalizability of the findings remains uncertain.

> We fully acknowledge the reviewer’s concern regarding data augmentation, and to address this issue, additional clarification has been provided in Section 4.1.2 (Data Preprocessing and Augmentation). Unlike natural images, top-down urban maps do not possess a fixed vertical orientation. Therefore, rotational augmentation of such spatial data does not merely produce duplicated samples; rather, it introduces diverse angular variations in street intersections and surrounding building configurations. As discussed by Shorten and Khoshgoftaar (2019), this type of geometric transformation helps mitigate positional and directional biases in CNN-based models and reduces overfitting to specific absolute orientations. In addition, the Conclusion (Section 6) has been revised to acknowledge the limitations of the current dataset and to propose dataset expansion as future work in order to further evaluate the generalization capability of the proposed model.

  1. The training strategy raises concerns about redundancy and potential overfitting. The rapid convergence and lack of validation on more heterogeneous and unseen contexts suggest that the model may be learning dataset-specific regularities rather than transferable context-response principles.

> We have addressed this concern as follows.

First, regarding validation on heterogeneous sites, the limitations of the dataset have been clarified in Section 4.1.1. The use of systematically planned new-town districts is explicitly justified as an intentional choice to enable the model to learn transferable architectural principles, rather than memorizing the morphological noise of irregular urban fabrics.

Second, the observed rapid convergence is interpreted as a characteristic of fine-tuning a pretrained diffusion model with additional spatial conditioning. As reported by Zhang et al. (2023, ICCV), models with strong pretrained backbones tend to adapt abruptly to the conditioning inputs—typically within fewer than 10,000 training steps—rather than gradually. In our study, the model reached its optimal performance at Epoch 8 (approximately 6,900 steps), which is consistent with these findings. This discussion has been added to Section 4.4.3.

In addition, to mitigate the risk of overfitting, an early stopping strategy was applied as described in Section 4.4.3. Training was terminated after 13 epochs based on validation loss monitoring, and the model weights from Epoch 8—corresponding to the best validation performance—were selected for subsequent analysis.

  1. The experimental design is too narrow, relying on a single ablation between a four-channel model and a boundary-only baseline. The absence of intermediate configurations and comparisons with alternative methods prevents an assessment of the actual contribution of the proposed approach.

> We agree with the reviewer that exploring intermediate configurations and conducting comparisons with alternative models would meaningfully broaden the scope of the study. In this regard, we acknowledge that the current experimental design has certain limitations. However, we would like to briefly clarify the rationale behind the adopted comparison strategy. The single-channel, site-boundary-based model (Model B) was not arbitrarily selected; rather, it serves as a baseline representing prior approaches that rely primarily on parcel boundaries for footprint generation, such as ArchiGAN (Chaillou, 2019). The primary objective of this study was to examine how the incorporation of richer multi-channel urban context influences the generative outcomes in comparison to such boundary-based approaches. In response to the reviewer’s valuable comment, we have explicitly acknowledged this limitation in the Conclusion of the revised manuscript and highlighted that future work should explore intermediate configurations (e.g., road networks vs. adjacent buildings) and include broader comparative evaluations with alternative models.

  1. The evaluation framework is inadequate, as it relies on entropy. While useful for measuring variability, entropy does not capture architectural quality and urban performance, yet the manuscript draws conclusions that go beyond what this metric can support.

> We agree with the reviewer that entropy alone is insufficient to fully capture architectural validity. To address this limitation, an additional quantitative metric, Intersection over Union (IoU), has been introduced in Section 5.1.3. A trend analysis of the average IoU between the generated footprints and the ground-truth (GT) designs was conducted across different probability thresholds. As shown in Figure 8, the results indicate that the proposed multi-channel model not only reduces entropy but also tends to produce spatial configurations that are more closely aligned with the GT layouts.

  1. The results lack statistical and qualitative rigor. No statistical significance analysis is provided and the qualitative assessment is based on selected examples without a systematic and objective evaluation protocol, introducing potential bias.

> We agree with the reviewer’s concern and have revised the evaluation framework to improve its rigor. To address the lack of statistical analysis, an IoU (Intersection over Union)–based trend analysis has been introduced in Section 5.1.3 to provide a quantitative examination of geometric validity. In addition, to reduce potential bias in the qualitative assessment, three evaluation criteria grounded in architectural and urban morphology theories have been established in Section 5.2, and the case studies have been analyzed in a more systematic manner based on these criteria.

  1. Methodological choices are not justified, reproducibility is incomplete and the manuscript overextends its claims, particularly regarding sustainability and urban performance, which are not evaluated.

> We fully accept the reviewer’s comment. The manuscript has been carefully revised to moderate the level of our claims throughout. In both the Introduction and the Conclusion, we have replaced statements that could be interpreted as asserting that context-aware placement directly guarantees urban performance. Instead, we now frame the proposed approach as providing a “morphological foundation” and a “geometrically coherent starting point” to support the development of sustainable design, rather than as directly achieving sustainability or urban performance outcomes.

Thank you again for your time and valuable feedback.

Reviewer 3 Report

Comments and Suggestions for Authors

This is an academic paper exploring the application of generative artificial intelligence in building outline generation. The authors propose an input method incorporating four-channel environmental information (roads, sidewalks, adjacent buildings, and site boundaries), using the conditional inpainting mechanism of LDM to generate building outlines. Overall, it's well-written. The paper has a rigorous structure, and the theoretical background is presented in a novel way. The research methodology is clear, especially the use of "pixel-wise information entropy" as a quantitative indicator to measure the reduction of uncertainty in model generation, which is a clever and scientific design. Suggestions for revision are as follows:

1. Lines 337-342: Supporting literature should be added, or a brief comparative experiment/visualization should be provided to demonstrate that this merging operation does not cause confusion or loss of key spatial information.

2. Line 372: A clear explanation is needed as to why the specific value of 10% was chosen. Are there any relevant ablation experiments to support this, or which previous classic studies/literature were referenced? Please provide citations or a brief explanation.

3. Lines 402-407: Quantitative validation relies almost entirely on "average information entropy" to demonstrate a reduction in model uncertainty. Low information entropy only indicates that the model-generated contours are more "focused" or "definite," but it does not directly prove that the generated contours are architecturally "excellent" or "reasonable." It is recommended to add at least one quantitative indicator that reflects architectural or geometric validity.

Author Response

Dear Reviewer,

We would like to thank you for your careful review and thoughtful comments on this manuscript. Please note that changes made to reflect the opinions of you and other reviewers are highlighted in red font in the revised manuscript. The changes made to reflect your opinion are as follows.

  1. Lines 337-342: Supporting literature should be added, or a brief comparative experiment/visualization should be provided to demonstrate that this merging operation does not cause confusion or loss of key spatial information.

> We fully acknowledge the reviewer’s concern regarding this issue. To address it, we conducted a VAE encoding–decoding reconstruction experiment and have added the results as Figure 5 (Spatial Information Preservation in Channel Merging with VAE) in Section 4.2.1. The visualization shows that the thin and hollow outlines of site boundaries and the filled polygonal forms of adjacent buildings remain distinguishable without spatial confusion or information loss, indicating that their geometric characteristics are preserved after reconstruction in the latent space.

  1. Line 372: A clear explanation is needed as to why the specific value of 10% was chosen. Are there any relevant ablation experiments to support this, or which previous classic studies/literature were referenced? Please provide citations or a brief explanation.

> We have clarified the rationale for the selected dropout rate in Section 4.2.3. The 10% dropout rate is now described as being motivated by the CFG study of Ho and Salimans (2022), and the relevant reference has been added to support this choice.

  1. Lines 402-407: Quantitative validation relies almost entirely on "average information entropy" to demonstrate a reduction in model uncertainty. Low information entropy only indicates that the model-generated contours are more "focused" or "definite," but it does not directly prove that the generated contours are architecturally "excellent" or "reasonable." It is recommended to add at least one quantitative indicator that reflects architectural or geometric validity.

> In response to the reviewer’s suggestion, a new section (Section 5.1.3: Geometric Validity via Intersection over Union) has been added. An IoU analysis was conducted between the generated footprint heatmaps and the ground-truth (GT) designs, and the results are presented in Figure 8. The analysis indicates that the proposed model not only concentrates the generated results but also tends to converge toward geometrically valid locations that are more consistent with plausible architectural layouts.

Thank you again for your time and valuable feedback.

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

All the comments have been addressed.

Reviewer 2 Report

Comments and Suggestions for Authors

For the authors’ benefit, I would only encourage a final proofread before production, particularly to ensure consistency in figure numbering, caption formatting and minor typographic presentation. These are editorial refinements and do not affect my recommendation for acceptance.

Back to TopTop