Next Article in Journal
Integration of Sustainable Urban Drainage Systems (SUDSs) in Highway Projects in Small Island Developing States (SIDSs) for Improved Resilience to Flooding
Previous Article in Journal
Peak-Hour Traffic Congestion and Level of Service Assessment Along the Jogeshwari-Vikhroli Link Road Corridor, Mumbai, India
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Proceeding Paper

Controlled GPT Augmentation of Landslide Inventories for Urban Resilience Analytics †

1
COEUS Institute, New Market, VA 22844, USA
2
Center for Trade and Investment, University of Dhaka, Dhaka 1000, Bangladesh
Presented at the 1st International Online Conference on Urban Sciences (IOCUS 2026), 20–22 May 2026; Available online: https://sciforum.net/event/IOCUS2026.
Environ. Earth Sci. Proc. 2026, 45(1), 13; https://doi.org/10.3390/eesp2026045013
Published: 31 August 2026

Abstract

Landslide inventories are essential for susceptibility modelling, risk communication, and urban resilience planning, yet many regional datasets remain incomplete, inconsistent, or weakly structured. This study presents a controlled GPT augmentation framework for improving the analytical usability of landslide inventories without replacing observed hazard records with synthetic events. The method was applied to 730 landslide records from the Chittagong Hill Area inventory. The original records were retained as the factual layer, while GPT generated only auxiliary annotations, including standardized triggers, uncertainty labels, exposure descriptors, severity labels, and process narratives. Baseline auditing showed substantial incompleteness, with missingness reaching 47.9% for water condition, 39.2% for distributional descriptor, 37.3% for material, and 31.2% for trigger information. The augmentation preserved 100% of observed empirical fields and reduced 13 nonblank raw trigger expressions into 8 standardized classes. All 730 records received structured process descriptors, and 274 retained moderate or high uncertainty labels. A comparative model-readiness test using fatality occurrence showed that augmented input retained 730 usable records and 33 fatal cases, compared with 375 records and 8 fatal cases in the complete-case original inventory. Extra Trees PR AUC increased from 0.190 to 0.609, and F1 from 0.187 to 0.427.

1. Introduction

Landslide inventories are central to disaster risk science because they provide the empirical basis for susceptibility mapping, hazard classification, casualty analysis, early warning, and urban resilience planning. Recent studies have increasingly used machine learning and artificial intelligence to analyse landslide and disaster inventories, including the application of AI to casualty analysis in the Chittagong region [1] and broader machine learning approaches for landslide susceptibility modelling [2,3]. These methods are attractive because landslides emerge from complex interactions among rainfall, slope morphology, soil and geological conditions, vegetation disturbance, road cutting, hill cutting, settlement exposure, and other human induced changes. However, the predictive strength of machine learning and deep learning models depends strongly on the quality of the inventory used for training. If the underlying records are incomplete, inconsistent, weakly structured, or affected by reporting errors, even advanced algorithms may produce unstable or poorly generalizable outputs.
This limitation is particularly important for regional landslide inventories, where the observed data may be valuable but not always complete enough for reliable computational modelling. The Chittagong Hilly Areas inventory, for example, provides a rare and important record of landslide events from 2001 to 2017, including information on failure type, damage, water condition, material, trigger, and settlement impact [4]. Yet, as with many disaster inventories, several explanatory fields may be missing, semantically ambiguous, or inconsistently encoded. Prior studies have shown that landslide susceptibility models are sensitive to inventory incompleteness, uncertain non landslide samples, class imbalance, and weak data representation [5]. For this reason, data augmentation has become an important direction in landslide modelling. Some studies have used generative adversarial networks to address class imbalance and improve spatial prediction [6,7]. These approaches demonstrate the value of augmentation, but they also raise a methodological question: how can additional data support machine learning without detaching the model from the empirical structure of observed landslide records?
Recent GPT-based geoscience research has shown that large language models can generate structured geoscience records and assist scientific data workflows [8]. However, fully synthetic landslide generation creates a risk of circular modelling if machine generated data are then used to train another machine learning system as if they were observed hazard evidence. This concern is strengthened by the well documented tendency of natural language generation systems to produce plausible but unsupported content [9]. The present study therefore proposes a more conservative framework: controlled GPT augmentation of observed landslide inventories. Instead of generating an entire landslide dataset from scratch, the method retains the original statistical, spatial, and categorical cues of the observed inventory while using GPT only to enrich incomplete descriptors, normalize inconsistent trigger labels, assign uncertainty indicators, and generate concise process narratives. In this design, the observed records remain the factual layer, and GPT functions as a constrained augmentation mechanism that improves modelling readiness without replacing empirical evidence. The aim is to make incomplete landslide inventories more usable for machine learning, deep learning, risk communication, and urban resilience analytics while avoiding the epistemic weakness of treating synthetic hazards as real observations.

2. Materials and Methods

Let the observed Chittagong Hill Area landslide inventory be denoted by
D O = { x i } i = 1 N , N = 730 ,
where each record x i contains empirical attributes including district, date, failure type, death count, settlement impact, area, material, water condition, damage intensity, economic impact, state, distributional descriptor, and trigger condition [4]. The original inventory was treated as the factual layer of the study. Therefore, all observed variables were preserved without modification, and GPT was restricted to generating auxiliary annotations rather than new landslide observations.
The inventory covers five districts, namely Chittagong, Rangamati, Coxś Bazar, Bandarban, and Khagrachari. It contains complete core empirical fields, including district, failure type, death count, and settlement impact, alongside incomplete explanatory fields such as water condition, material, damage intensity, economic impact, distributional descriptor, and trigger condition. This combination made the dataset suitable for testing whether GPT-assisted annotation could improve semantic completeness while retaining the observed statistical and spatial structure of the landslide records.
For each variable v j , missingness was quantified as
M j = 1 N i = 1 N I ( x i j { , blank , NA } ) × 100 ,
where x i j denotes the value of variable v j in record x i , and I ( · ) is the indicator function. This audit identified fields with high incompleteness and guided the augmentation targets. Before GPT augmentation, deterministic normalization was applied to lexical inconsistencies in categorical fields. In particular, raw trigger expressions were mapped into a controlled taxonomy
ϕ : T raw T std ,
where T raw represents the set of original trigger labels and T std represents standardized trigger classes.
GPT-based augmentation was then applied at the row level. For each observed record x i , the model generated an annotation vector
a i = g θ ( x i , C ) ,
where g θ denotes the constrained GPT function and C denotes the augmentation constraints. These constraints prohibited alteration of observed fields, creation of new landslide events, and confident completion where empirical evidence was insufficient. The annotation vector was defined as
a i = ( t i , c i , s i , e i , u i , r i , p i ) ,
where t i is the standardized trigger class, c i is trigger confidence, s i is severity class, e i is exposure context, u i is uncertainty label, r i is the process narrative, and p i is the planning summary. The augmented dataset was therefore expressed as
D A = { ( x i , a i ) } i = 1 N ,
so that empirical records and generated annotations remained analytically separable.
Validation was performed using preservation, consistency, and uncertainty checks. Field preservation was assessed as
P = 1 N i = 1 N I ( x i O = x i A ) ,
where x i O and x i A denote the original and post augmentation observed fields. A valid augmentation required P = 1 . Taxonomic compression was measured by comparing the number of raw and standardized trigger categories,
R T = 1 | T std | | T raw | ,
where larger values indicate stronger reduction of redundant categorical noise. Finally, records with weak empirical support were assigned explicit uncertainty labels rather than imputed as factual values. This ensured that augmentation improved semantic completeness while preserving the evidential boundary between observed landslide records and GPT generated explanatory metadata.
Domain validation was added through an independent review by Prof. Edris Alam of Business Continuity Management & Integrated Emergency Management, Rabdan Academy, Abu Dhabi, UAE. The review focused on whether the GPT-augmented trigger classes, uncertainty labels, exposure descriptions, and process narratives were plausible and consistent with the observed landslide attributes.
To test model readiness, a comparative classification experiment was also conducted using fatality occurrence as the observed target, where y i = 1 if Death i > 0 , and y i = 0 otherwise. The original inventory was evaluated under complete predictor availability, while the augmented inventory used the same observed fields with structured GPT-generated annotations. To avoid target leakage, augmented severity class, process narrative, and planning summary were excluded from the predictors. Logistic regression, random forest, and extra trees classifiers were evaluated using stratified five-fold cross-validation.
The complete code used for missingness auditing, trigger standardization, GPT-assisted annotation using OpenAI GPT-3.5-turbo, validation, and figure generation is publicly available at https://github.com/DrSufi/GPT-Augmented-Landslide (accessed on 31 May 2026).

3. Results

The controlled augmentation experiment was conducted on 730 observed landslide records from the Chittagong Hill Area inventory. The initial audit of the dataset showed that the inventory was empirically valuable but analytically incomplete. As reported in Table 1, several explanatory variables had substantial missing or blank values. The highest incompleteness was observed for water condition, with 350 missing records, representing 47.9% of the dataset. This was followed by distributional movement descriptor at 39.2%, material type at 37.3%, and damage intensity, economic impact, and primary damage indicator at approximately 37.1% each. Trigger information was also absent or blank in 228 records, equivalent to 31.2% of the dataset. In contrast, core empirical fields such as district, failure type, death count, and settlement impact were complete. This pattern suggests that the dataset is sufficiently reliable for empirical anchoring, but requires augmentation to improve explanatory depth, semantic consistency, and interpretability.
The spatial distribution of records showed clear district level variation. As summarized in Table 2, Chittagong contained the highest number of landslide records, with 208 events and 146 recorded deaths. Rangamati followed with 193 records but only 14 deaths, while Cox’s Bazar had 124 records and 26 deaths. Bandarban and Khagrachari contained 118 and 87 records respectively, with lower mortality totals. The district level results indicate that event frequency and human impact are not evenly aligned. Chittagong represents the strongest mortality burden, whereas Rangamati and Cox’s Bazar show substantial event presence with different severity profiles. This distinction is important because it shows that landslide risk cannot be interpreted from event counts alone.
The failure type distribution further demonstrates the structural heterogeneity of the inventory. Table 3 shows that slide was the most frequent failure type, accounting for 285 records and 113 deaths. Flow was the second most frequent category, with 230 records and 26 deaths. Fall, unrecognized, complex, and topple events occurred less frequently, although the mortality pattern was not proportional to frequency. Topple events, for example, appeared in only 17 records but were associated with 33 deaths, suggesting a comparatively high fatality burden relative to their occurrence. Figure 1 provides a clearer view of this district by failure type concentration. Chittagong was dominated by slide events, Rangamati and Bandarban showed stronger flow concentrations, and Cox’s Bazar included a notable number of unrecognized landslides. These results confirm that any augmentation process must preserve the empirical structure of the observed inventory rather than imposing a uniform synthetic pattern across all districts.
The trigger field provided the clearest demonstration of the value of controlled augmentation. The raw data contained inconsistent trigger descriptions, including spelling variations and overlapping expressions related to rainfall, hill cutting, road cutting, construction, agriculture, and load induced instability. Through the augmentation process, these raw descriptions were normalized into eight interpretable trigger classes, as shown in Table 4. Rainfall was the dominant standardized trigger, accounting for 390 records, or 53.4% of the dataset, and was associated with 157 deaths. Rainfall combined with hill cutting was the second most important trigger class, with 96 records and 32 deaths. A further 228 records remained classified as unknown because the original trigger field was missing or insufficiently specified. This is an important outcome because the augmentation did not conceal uncertainty. Instead, uncertain records were retained as unknown, thereby preserving analytical caution.
Figure 2 extends this result by showing the relationship between district, standardized trigger class, record count, and mortality. The bubble pattern shows that rainfall-related landslides dominate across the inventory, but their human impact is particularly concentrated in Chittagong. The figure also shows that compound triggers such as rainfall and hill cutting are especially visible in Cox’s Bazar and other districts, indicating the relevance of both climatic and human modified slope conditions. The inclusion of death labels within the bubbles makes the plot useful for resilience interpretation, because it separates high frequency trigger patterns from high impact trigger patterns.
The validation summary in Table 5 further supports the methodological reliability of the augmentation process. The observed empirical fields were preserved completely, meaning that the augmentation layer did not overwrite the original landslide records. The raw trigger field contained 13 nonblank categories, many of which reflected spelling variants or overlapping descriptions. These were reduced into eight standardized classes, improving interpretability while retaining an explicit unknown class for unresolved records. All 730 records received a structured process descriptor, while 274 records were assigned moderate or high uncertainty. This outcome is methodologically important because the augmentation process did not simply fill missing data with confident assumptions. Instead, it enriched records where possible and retained uncertainty where the original evidence was insufficient. The independent expert review further indicated that the augmented annotations were broadly consistent with the observed landslide attributes and suitable for use as auxiliary metadata rather than as new empirical observations.
To empirically assess model readiness, fatality occurrence was used as a binary target and three classifiers were compared under original and augmented feature settings. As shown in Table 6, the augmented inventory retained all 730 records and all 33 fatality cases, whereas the complete-case original inventory retained 375 records and 8 fatality cases. Predictive metrics also improved across the three classifiers. For logistic regression, ROC AUC increased from 0.897 to 0.918 and PR AUC increased from 0.485 to 0.575. The extra trees model showed the largest gain, with PR AUC increasing from 0.190 to 0.609 and F1 increasing from 0.187 to 0.427.
Overall, the results show that the proposed augmentation strategy improves the analytical usability of the CHA landslide inventory without altering its empirical foundation. The original records remain the factual layer, while the augmentation process adds standardized trigger classes, severity labels, exposure interpretations, uncertainty indicators, and short process narratives. The main contribution of the experiment is therefore not the creation of artificial landslide evidence, but the transformation of a sparse and unevenly documented inventory into a more structured and interpretable dataset. This makes the dataset more suitable for urban resilience analytics, risk communication, and future modelling, while maintaining a clear distinction between observed landslide records and generated explanatory annotations. The GPT-augmented landslide dataset generated in this study, together with the code used to generate the figures and comparative model-readiness experiment, is publicly available at https://github.com/DrSufi/GPT-Augmented-Landslide (accessed on 21 July 2026).

4. Discussion and Concluding Remarks

The findings indicate that GPT can be useful for landslide inventory augmentation when its role is narrowly defined, empirically constrained, and independently checked. In this study, the Chittagong Hill Area records were kept as the factual basis, while GPT was used only to add auxiliary information, including standardized trigger labels, uncertainty indicators, exposure descriptions, and short process narratives. The added expert review supports the plausibility of these annotations as auxiliary metadata, and the comparative machine learning test provides direct evidence of improved pre-modelling data readiness. This distinction matters because landslide susceptibility and impact models depend heavily on the quality and completeness of the inventory used for training [2,3,5]. Compared with approaches that create additional synthetic samples for class balancing or spatial prediction [6,7], the present framework is more cautious because it preserves observed records, makes uncertainty explicit, and improves the interpretability of fields that were originally sparse or inconsistently described. Operationally, such annotations could support hazard assessment and urban resilience planning by helping analysts screen incomplete records, identify dominant trigger contexts, separate observed evidence from uncertain interpretation, and prepare more consistent inputs for susceptibility mapping, local risk communication, and field verification.
The study remains limited by its use of a single regional inventory from the Chittagong Hill Area, and the results may not transfer directly to inventories produced under different geomorphological, climatic, institutional, or reporting conditions. The comparative model-readiness experiment used fatality occurrence as a compact observed target and should not be interpreted as a full operational landslide susceptibility benchmark. The generated annotations also depend on the prompt, available input fields, and model behaviour, which remains important given the tendency of language models to produce plausible but unsupported outputs [9]. Future work should test the framework on additional landslide inventories, include larger expert validation samples, compare different GPT models and prompt designs, and evaluate whether the enriched fields improve susceptibility classification, severity estimation, or cross-regional transferability when combined with rainfall, land use, remote sensing, and geotechnical variables.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

For reproducibility and validation, the GPT-augmented dataset and code used in this study are publicly available at https://github.com/DrSufi/GPT-Augmented-Landslide (accessed on 21 July 2026).

Acknowledgments

The author would like to thank Edris Alam of Business Continuity Management & Integrated Emergency Management, Rabdan Academy, Abu Dhabi, UAE, for validating the GPT-augmented dataset.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Alam, E.; Sufi, F.; Islam, A.R.M.T. A scenario-based case study: Using AI to analyze casualties from landslides in Chittagong Metropolitan Area, Bangladesh. Sustainability 2023, 15, 4647. [Google Scholar] [CrossRef] [Scilit]
  2. Youssef, A.M.; Pourghasemi, H.R. Landslide susceptibility mapping using machine learning algorithms and comparison of their performance at Abha Basin, Asir Region, Saudi Arabia. Geosci. Front. 2021, 12, 639–655. [Google Scholar] [CrossRef] [Scilit]
  3. Ado, M.; Amitab, K.; Matori, A.R.; Yusof, K.N.A.M.; Pradhan, B. Landslide susceptibility mapping using machine learning: A literature survey. Remote Sens. 2022, 14, 3029. [Google Scholar] [CrossRef] [Scilit]
  4. Rabby, Y.W.; Li, Y. Landslide inventory (2001–2017) of Chittagong Hilly Areas, Bangladesh. Data 2020, 5, 4. [Google Scholar] [CrossRef] [Scilit]
  5. Huang, F.; Mao, D.; Jiang, S.H.; Zhou, C.; Fan, X.; Zeng, Z.; Catani, F.; Yu, C.; Chang, Z.; Huang, J.; et al. Uncertainties in landslide susceptibility prediction modeling: A review on the incompleteness of landslide inventory and its influence rules. Geosci. Front. 2024, 15, 101886. [Google Scholar] [CrossRef] [Scilit]
  6. Al-Najjar, H.A.H.; Pradhan, B.; Sarkar, R.; Beydoun, G.; Alamri, A. A new integrated approach for landslide data balancing and spatial prediction based on generative adversarial networks (GAN). Remote Sens. 2021, 13, 4011. [Google Scholar] [CrossRef] [Scilit]
  7. Al-Najjar, H.A.; Pradhan, B. Spatial landslide susceptibility assessment using machine learning techniques assisted by additional data created with generative adversarial networks. Geosci. Front. 2021, 12, 625–637. [Google Scholar] [CrossRef] [Scilit]
  8. Sufi, F.K. A framework for integrating GPT into geoscience research. J. Econ. Technol. 2026, 4, 226–237. [Google Scholar] [CrossRef] [Scilit]
  9. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.; Chen, D.; Dai, W.; et al. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef] [Scilit]
Figure 1. District level concentration matrix of observed landslide failure types.
Figure 1. District level concentration matrix of observed landslide failure types.
Eesp 45 00013 g001
Figure 2. Augmented trigger taxonomy by district. Bubble size represents record count and internal labels indicate total deaths. Bubble color is used only for visual distinction and does not encode an additional variable.
Figure 2. Augmented trigger taxonomy by district. Bubble size represents record count and internal labels indicate total deaths. Bubble color is used only for visual distinction and does not encode an additional variable.
Eesp 45 00013 g002
Table 1. Baseline missingness audit of the observed CHA landslide inventory.
Table 1. Baseline missingness audit of the observed CHA landslide inventory.
FieldMissing RecordsMissing Percentage
Water_Cont35047.9%
Distri_28639.2%
Material27237.3%
Damae_Int227137.1%
Economic27137.1%
Dam_Int127137.1%
State23131.6%
Location22931.4%
Triggers_22831.2%
Area7810.7%
Date81.1%
District00.0%
Fail_Type00.0%
Death_00.0%
Settlemet_00.0%
Table 2. District level empirical profile after augmentation.
Table 2. District level empirical profile after augmentation.
DistrictRecordsDeathsFatal RecordsMean AreaHigh or Extreme
Chittagong20814615780.820
Rangamati1931431892.635
Cox’s Bazar12426102778.821
Bandarban1181241804.811
Khagrachari8721132.811
Table 3. Failure type profile of the observed landslide inventory.
Table 3. Failure type profile of the observed landslide inventory.
Failure TypeRecordsDeathsMean AreaHigh or Extreme
Slide285113910.730
Flow23026773.438
Fall8711334.312
Unrecognized7764989.56
Complex34113913.27
Topple1733678.85
Table 4. Trigger standardization after controlled augmentation.
Table 4. Trigger standardization after controlled augmentation.
Augmented Trigger ClassRecordsShareDeathsConfidence Pattern
Rainfall39053.4%157High
Unknown22831.2%0Low
Rainfall and hill cutting9613.2%32High
Rainfall and road cutting91.2%5High
Rainfall and slope agriculture40.5%0High
Hill cutting10.1%0Medium
Rainfall and construction activity10.1%0High
Load induced instability10.1%6Medium
Table 5. Validation summary of the controlled augmentation process.
Table 5. Validation summary of the controlled augmentation process.
Validation ItemResultInterpretation
Observed field preservation100%Original empirical fields were not overwritten
Original trigger categories13 nonblank categoriesRaw trigger field contained spelling and naming variation
Augmented trigger categories8 standardized classesTrigger noise was reduced into an interpretable taxonomy
Rows with semantic descriptor730Every record received a structured process narrative
Rows with moderate or high uncertainty274Uncertainty was retained rather than hidden
Independent expert reviewCompletedAugmented annotations were judged plausible as auxiliary metadata
Table 6. Comparative machine learning assessment of model readiness using fatality occurrence as the observed target.
Table 6. Comparative machine learning assessment of model readiness using fatality occurrence as the observed target.
ModelFeaturesRecordsFatalBal. Acc.ROC AUCPR AUCF1
Logistic reg.Original37580.7300.8970.4850.330
Logistic reg.Augmented730330.7930.9180.5750.364
Random forestOriginal37580.5880.8630.3500.150
Random forestAugmented730330.7730.9090.5540.393
Extra treesOriginal37580.6280.8780.1900.187
Extra treesAugmented730330.7930.9160.6090.427
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sufi, F. Controlled GPT Augmentation of Landslide Inventories for Urban Resilience Analytics. Environ. Earth Sci. Proc. 2026, 45, 13. https://doi.org/10.3390/eesp2026045013

AMA Style

Sufi F. Controlled GPT Augmentation of Landslide Inventories for Urban Resilience Analytics. Environmental and Earth Sciences Proceedings. 2026; 45(1):13. https://doi.org/10.3390/eesp2026045013

Chicago/Turabian Style

Sufi, Fahim. 2026. "Controlled GPT Augmentation of Landslide Inventories for Urban Resilience Analytics" Environmental and Earth Sciences Proceedings 45, no. 1: 13. https://doi.org/10.3390/eesp2026045013

APA Style

Sufi, F. (2026). Controlled GPT Augmentation of Landslide Inventories for Urban Resilience Analytics. Environmental and Earth Sciences Proceedings, 45(1), 13. https://doi.org/10.3390/eesp2026045013

Article Metrics

Back to TopTop