Next Article in Journal
Integrated Assessment of Groundwater Quality, Infrastructure, and Livestock-Watering Suitability in Western Kazakhstan
Previous Article in Journal
Recycled Concrete Aggregates in Bioretention Cells: Pollutant Retention, Start-Up and Stabilization Dynamics, and Resource Conservation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Beyond Composite Suitability: Evidence-Aware Pareto Frontiers for Geothermal Exploration Decisions in Colombia

by
Carlos Robles-Algarín
1,*,
Víctor Olivero-Ortiz
1,* and
Carolina Diosa Rosas
2
1
Facultad de Ingeniería, Universidad del Magdalena, Santa Marta 470004, Colombia
2
Agencia Nacional de Hidrocarburos (ANH), Bogotá 111321, Colombia
*
Authors to whom correspondence should be addressed.
Resources 2026, 15(9), 119; https://doi.org/10.3390/resources15090119
Submission received: 14 August 2026 / Revised: 5 September 2026 / Accepted: 10 September 2026 / Published: 14 September 2026

Abstract

Regional geothermal screening often collapses prospectivity, territorial constraints, institutional support, and data maturity into a compensatory score, obscuring why an area ranks highly. We develop an evidence-aware, non-compensatory framework for Colombia. Eleven criteria were standardized at 300 m and weighted using Best–Worst Method judgments from 15 experts. Results for 23 registered areas were decomposed into geoscientific prospectivity, territorial compatibility, institutional alignment, and evidence gap. Institutional alignment remained a contextual implementation dimension but was excluded from Pareto dominance. Paired frontiers distinguish advancement screening from information-acquisition screening. Robustness was evaluated through 10,000 expert-panel bootstrap replicates, 81 transformations of four dominant geoscientific layers, and directed missing-data tests. Cubarral, Nereidas, Iza, Paipa, and Santa Rosa formed the advancement frontier, which was unchanged across all transformation scenarios; joint robustness frequencies ranged from 83.4% to 100%. However, neutral score-3 imputation retained only Cubarral, Nereidas, and Santa Rosa (Jaccard = 0.600). Eight of nine learning-front members persisted across all transformation scenarios, while Santa Rosa was transformation-sensitive. Apiay ranked second in every expert bootstrap but belonged to neither frontier, showing that rank stability does not imply non-dominance. The framework separates preference uncertainty from evidence-model and missing-data assumptions while providing screening sets requiring costed feasibility appraisal.

1. Introduction

Geothermal resources can provide non-variable renewable heat and electricity, yet their early development is constrained by a distinctive asymmetry: High-cost decisions are often required before the resource is fully characterized [1]. National and regional assessments are therefore expected to do more than indicate where temperature anomalies or surface manifestations occur [2,3]. They must also help decision makers determine where additional geochemistry, geophysics, slim hole drilling, community engagement, and environmental studies can most effectively reduce uncertainty. In data-limited settings, this decision is particularly difficult because the absence of observations may reflect limited exploration rather than the absence of a geothermal resource [1,2]. Recent depth-integrated fuzzy MCDA also demonstrates how surface and subsurface evidence can be combined and evaluated at a controlled geothermal test site [4].
Geographic information systems (GISs) and multicriteria decision analysis (MCDA) have become established tools for integrating geological, geochemical, hydrogeological, remote sensing, and accessibility evidence into geothermal favorability assessments [4,5,6,7,8,9,10,11]. Applications span regional geothermal exploration, resource potential assessment, play fairway analysis, and screening for direct use opportunities [5,6,7,8,9,10,11]. Knowledge-driven approaches are particularly useful when consistent regional or national subsurface data are unavailable [5,6,7,8], while geothermal play-fairway approaches provide a framework for organizing multiple lines of evidence and their associated confidence [9]. Nevertheless, many regional applications remain centered on a single weighted or favorability surface [5,6,7,8,10,11]. Weighted aggregation is transparent and operationally convenient, but it is also compensatory because a strong score in one criterion can offset weak evidence or an unfavorable territorial condition in another. This property becomes problematic when a screening result is interpreted as an indication of uniform project readiness.
Recent geothermal play-fairway studies make the distinction between favorability and confidence increasingly explicit. The Snake River Plain workflow organized geologic, geophysical, and hydrologic evidence into resource-risk products with accompanying confidence information [12]; comparable applications in the Modoc Plateau and Hawai‘i combined evidence, uncertainty, resource probability, and development-viability screening to support staged exploration [13,14,15]. Bayesian evidential-learning and value-of-information research addresses a later decision stage by quantifying how new observations may change predicted resource outcomes [16]. Together, these studies show that uncertainty is decision-relevant rather than merely a cartographic qualifier. However, most regional workflows still culminate in a single favorability or risk product, whereas national agencies must decide both which areas to advance and where additional information has the greatest strategic value.
Colombia provides a relevant case for examining this problem. Its geothermal systems are associated mainly with the Andean volcanic and tectonic setting, with identified hydrothermal areas in the Central Cordillera and localized sectors of the Eastern and Western Cordilleras. The national volumetric assessment considered 324 hot springs, grouped into 165 clusters, and estimated approximately 1170.2 MWe across 21 preliminary geothermal areas [17]. However, the progression from regional resource assessment to exploration investment remains uneven. This setting provides an opportunity to examine whether favorable composite conditions necessarily correspond to exploration readiness and whether areas with unresolved evidence may instead warrant priority for additional information acquisition.
To address this limitation, this study moves beyond a single composite suitability score by separating geoscientific prospectivity, territorial compatibility, institutional alignment, and evidence gaps. Institutional alignment is retained as a contextual implementation dimension but does not enter Pareto dominance, which is constructed from prospectivity, compatibility, and evidence gap. A paired Pareto framework is then used to identify areas that are non-dominated either for advancement or for information acquisition, while robustness is evaluated under uncertainty in expert preferences and alternative transformations of key geoscientific evidence. This design allows stable composite rankings to be contrasted with non-compensatory decision efficiency and distinguishes areas that are robustly attractive for advancement from those for which their principal screening rationale is to reduce consequential uncertainty.

Research Gap and Questions

The unresolved problem is not how to add another geothermal suitability map to the literature. It is how to prevent a regional weighted score from being mistaken for uniform exploration readiness. This study addresses this problem through three questions:
  • How do registered Colombian geothermal areas differ when geoscientific prospectivity, territorial compatibility, institutional alignment, and evidence gaps are reported separately?
  • Which areas remain Pareto efficient for advancement or information acquisition under uncertainty in expert preferences and geoscientific evidence transformations?
  • How does non-compensatory efficiency change the interpretation of stable composite ranks and the official geothermal inventory?
The proposed answer is a paired frontier analysis for which its stability is evaluated under both expert preference and geoscientific transformation uncertainty. The advancement frontier identifies areas that are not dominated when low evidence risk is desirable, whereas the learning frontier identifies areas that are not dominated when unresolved evidence is treated as an opportunity to reduce consequential uncertainty. Secondary action labels support communication but do not replace continuous indicators, project scale validation, or reserves classification.

2. Materials and Methods

2.1. Study Context and Units of Analysis

This study covered continental Colombia using MAGNA-SIRGAS 2018/Origen-Nacional (EPSG:9377). Two connected units were analyzed. The first was a 300 m raster screening envelope defined by the availability of the transformed geothermal evidence layers. The second comprised 23 environmentally corrected polygons associated with geothermal areas or registrations. These polygons were used to summarize scores, uncertainty, evidence gaps, estimated power potential, and environmental area retention. Their interpretation is regional and exploratory: polygon means do not represent a confirmed reservoir, productive well location, or bankable capacity.
The geothermal areas correspond to recognized Andean prospects and additional registrations associated with hydrocarbon operations. The database includes an estimated electric potential attribute (MWe), development-status descriptions, environmental-conflict attributes, original polygon areas, and corrected geometries. For national contexts, municipal and departmental boundaries, official geothermal-area polygons, aquifer information, and thematic rasters were available in the project geodatabase. The project geodatabase contains 326 thermal-spring records representing 324 distinct mapped locations. Although the 2021 national volumetric assessment also considered 324 springs [17], the two counts refer to different data structures: in the project geodatabase, two pairs of separately identified springs share mapped coordinates.

2.2. Criteria, Transformations, and Semantic Audit

Eleven conceptual criteria were retained from the expert-elicitation design (Table 1). Continuous and categorical source variables had previously been transformed to a common 1–5 scale, where 5 represents the most favorable implemented condition and 1 represents the least favorable condition. Thermal springs and interpolated fluid temperature were evaluated within a 5 km exploration influence used by the project. The geothermal gradient combined the national bottom-hole-temperature surface [18] with geothermal-area-specific information where available. Estimated power potential followed the national volumetric resource inventory [17] based on the regional volumetric assessment framework [3]. Aquifer evidence was based on the national aquifer systems map [19]. Environmental sensitivity incorporated RUNAP protected areas, Ramsar wetlands, and páramos. The governmental-interest criterion was derived from a documentary review of 27 departmental and 226 municipal development plans. The 5 km distance was fixed a priori as a reconnaissance support radius because mapped manifestations may be displaced by several kilometers from productive or upflow zones and comparable direct-use screening has used a 5 km influence envelope [6,11]. The temperature interpolation was restricted to the same radius to limit extrapolation beyond the local observation support. Supplementary File S5 reports the original variable, source, and year available in the retained project documentation; transformation thresholds; categorical coding; missing-value handling; and spatial preprocessing for every criterion. Where source-layer snapshot dates were not retained in the analytical export, S5 states that limitation explicitly.
A semantic and structural audit preceded re-integration. Departmental and municipal governmental-interest layers were averaged so that a criterion with two administrative representations did not receive twice its elicited weight. Mining-use and general land-use representations were likewise averaged into one land-compatibility criterion. National and geothermal-area gradient surfaces were treated as alternative spatial coverage for the same criterion. Territorial diversification was assigned 5 inside identified geothermal areas and 1 outside them. Finally, although the project report named the last criterion ‘scarcity of detailed studies’, its implemented 1–5 scale assigned 5 to the presence of exploratory geothermal or hydrocarbon wells and 1 to their absence. The present analysis therefore uses the monotonic and auditable name ‘detailed-study evidence’ or ‘evidence maturity’. This correction changes only the interpretation of the existing scale, not the underlying cell values. For governmental interest, each departmental and municipal plan received 5 when it contained an explicit geothermal or non-excluding FNCER project, indicator, or target, and it received 1 otherwise; averaging the two administrative signals produced possible combined values of 1, 3, or 5. This variable measures documentary policy attention, not institutional capacity, budget execution, permitting performance, or durable political commitment.

2.3. Best–Worst Method and Group Aggregation

Fifteen experts evaluated the geothermal criteria using the linear Best–Worst Method (BWM) [20]. Expert judgments were analyzed and reported only in aggregate, without associating individual responses with identifiable participants. Participants authorized the processing of the information provided for technical and academic purposes. Each participant identified the most important criterion (Best) and the least important criterion (Worst) and completed the corresponding Best-to-Others and Others-to-Worst comparisons using the 1–9 preference scale. For expert k , the linear optimization model minimized the maximum deviation ξ k , subject to the Best-to-Criterion and Criterion-to-Worst constraints, non-negativity, and a unit-sum weight vector. Individual ξ k values ranged from 0.0672 to 0.1737 (median of 0.1029), providing a transparent indication of internal fit rather than a binary exclusion rule [21,22].
The panel was assembled through purposive, non-probability technical recruitment for geothermal-specific judgment rather than population representation. The questionnaire recorded current role, organization, energy-sector experience, and a free-text statement of geothermal experience; because it requested current position rather than academic degree, disciplinary breadth is described from self-reported roles without inferring unreported credentials. Affiliations comprised public energy, mining, environmental, or geoscience agencies (7/15; 46.7%); universities, research, or postgraduate training (4/15; 26.7%); industry, consulting, or utility organizations (3/15; 20.0%); and a professional association (1/15; 6.7%). Seven participants reported more than five years of energy-sector experience, one reported three to five years, and seven reported fewer than three years. Eleven participants (73.3%) described geothermal-related experience in research, exploration, project development, regulation, community relations, or sector organizations. Self-reported roles included geologists and geothermal specialists or consultants, academic researchers, environmental and regulatory professionals, community-relations personnel, and sector-support roles. All 15 completed BWM responses were retained. These characteristics document the breadth of the observed panel but do not make it representative of Colombia’s geothermal expert or stakeholder population.
The individual weight vector of each expert was normalized to sum to one. Let w j k denote the normalized weight assigned to criterion j by expert k . Group weights were obtained using the normalized geometric mean, which is appropriate for aggregating ratio-based judgments [23]. For K experts and J criteria, the aggregated weight w j of criterion j was calculated as follows:
w j = k = 1 K w j k 1 / K l = 1 J k = 1 K w l k 1 / K , j = 1 , , J
where K = 15 is the number of experts, J = 11 is the number of criteria, and w j k is the normalized weight assigned to criterion j by expert k .
The clean composite score S i for a cell or registered area i was then calculated as follows:
S i = j = 1 J w j x i j
where the standardized score and aggregated weight retain the notation defined above. For the national cellwise screening surface only, a missing transformed raster value at an otherwise valid screening location received the least favorable score of 1 so that the term could not disappear from the weighted sum. Registered-area criterion means and diagnostic dimensions were instead calculated over valid cells; missing cells were not assigned 1 in those area means. Their influence on the Pareto analysis was carried separately through criterion-specific coverage and the evidence-gap indicator. The two operations therefore apply to different outputs and do not impose a double penalty on area-level prospectivity.

2.4. Dimension-Separated Diagnostics

Four diagnostics were calculated for each registered area. Geoscientific prospectivity was the within-dimension weighted mean of estimated power potential, thermal springs, fluid temperature, geothermal gradient, and aquifer evidence. Territorial compatibility was the within-dimension weighted mean of sensitive ecosystems, land compatibility, and distance from Indigenous territories. Institutional alignment was the within-dimension weighted mean of governmental interest and territorial diversification. Detailed-study evidence was reported separately. All diagnostic scores retained the 1–5 scale. Within each diagnostic dimension, the relevant area-level criterion means were combined using BWM weights renormalized to sum to one within that dimension. No normalization or aggregation was then applied across the three Pareto objectives.
Technical-data coverage was calculated as the geoscientific-weighted proportion of polygon cells with valid evidence for each of the five prospectivity criteria. An evidence-gap indicator combined (i) the absence of detailed-study evidence and (ii) missing technical coverage. For registered area i , the indicator was defined as
G a p i α = 100 α 1 M i + 1 α 1 C i
where M i is the detailed-study evidence score rescaled from 1–5 to 0–1, C i is technical-data coverage expressed as a proportion, and α controls the relative contribution of evidence maturity and data coverage. The baseline analysis used α = 0.5 (equal weighting). Sensitivity cases used α = 0.3 and α = 0.7 , corresponding to 30/70 and 70/30 splits, respectively. The evidence gap approaches zero where both detailed-study evidence and spatial data coverage are strong. Both advancement and learning frontiers, not only the secondary labels, were recomputed for every alpha specification.
A directed missing-data sensitivity test nevertheless recomputed each area-level geoscientific criterion as c k s ¯ k + ( 1 c k ) m , where s ¯ k is the valid-cell mean, c k is criterion-specific coverage, and m is an imputed missing-cell score. Two alternatives were evaluated: conservative imputation ( m = 1 ) and neutral imputation ( m = 3 ). Geoscientific prospectivity and both Pareto frontiers were recomputed while retaining the original evidence-gap coverage term, thereby testing whether frontier identity depends on treating missing cells as part of prospectivity.

2.5. Non-Compensatory Advancement and Learning Frontiers

Two Pareto frontiers were calculated from the dimension-separated area scores. The advancement frontier maximized geoscientific prospectivity and territorial compatibility while minimizing evidence gap. It represents areas for which no other registered area was at least as favorable on all three dimensions and strictly more favorable on at least one. The learning frontier maximized prospectivity, compatibility, and evidence gap. A high gap is desirable only in this second objective because it represents the opportunity to reduce consequential uncertainty in an otherwise relevant area. Thus, the learning frontier is a preliminary information-gap screening tool, not a value-of-information calculation, cost-effectiveness analysis, positive recommendation, or funding rule. The learning frontier operates on registered areas as portfolio units and does not itself optimize adjacency, road or grid access, survey mobilization, or spatial clustering; membership is a pre-feasibility information-acquisition screening signal that requires a subsequent logistics and cost screen.
Pareto dominance was evaluated without assigning trade-off weights across the three dimensions. This prevents a large gain in one dimension from automatically compensating for a loss in another and distinguishes the method from a second weighted-sum ranking. Institutional alignment was reported as contextual information but excluded from dominance because an administrative signal should not determine geological or territorial efficiency. For each area, the analysis retained baseline frontier membership and the number of other areas that dominated it. This exclusion also prevents a short-term documentary policy signal from creating non-dominance independently of geoscientific prospectivity, territorial compatibility, and evidence status; institutional alignment therefore remains a contextual implementation variable. Dominance was evaluated using full-precision area diagnostics and a numerical tolerance of ε = 10 12 . An area was considered to dominate another when it was no worse than the other area, within ε , on all objectives and better by more than ε on at least one objective. The corresponding supplementary analysis script implements this rule exactly.

2.6. Expert-Panel Resampling and Frontier Stability

Sensitivity to the composition of the observed expert panel was evaluated through a non-parametric bootstrap. In each of 10,000 replicates, 15 experts were sampled with replacement from the observed panel. Group weights were recalculated through the geometric mean, applied to the 23 area-level criterion vectors, and converted to ranks. The diagnostic dimensions and evidence gap were then recomputed, and both Pareto frontiers were reevaluated. For every criterion and area, the analysis reports the 2.5th and 97.5th percentiles. Area-level outputs include the score interval, median rank, rank interval, empirical frequencies of appearing in the top three and top five, and empirical frequencies of membership in each frontier. Reporting how often an alternative retains a rank or decision status follows the broader acceptability principle used in decision analysis under uncertain preferences [24]; here, however, the sampling distribution is the observed expert panel rather than an unrestricted weight space. The procedure does not represent all sources of uncertainty, including measurement error and alternative transformations, but directly tests sensitivity to expert panel composition. The same resampled expert indices were applied jointly across all criteria, preserving each respondent’s complete preference vector and the within-expert dependence among criterion weights. The elicitation did not collect criterion-specific confidence or expertise scores, so the bootstrap does not estimate a hierarchical expertise model; criterion-weighted resampling is a priority for future elicitation designs. Accordingly, these percentiles and membership frequencies describe perturbations of the finite observed panel and are not population confidence intervals or estimates of population-level expert uncertainty.

2.7. Geoscientific-Transformation Sensitivity

Transformation-model uncertainty was evaluated for estimated MWe, thermal springs, fluid temperature, and geothermal gradient. These four layers account for 0.4333 of the total BWM weight and 84.35% of the geoscientific dimension. For every valid 300 m cell in the original geodatabase, the delivered standardized score s was first mapped to the unit interval as
u = s 1 4
and then transformed according to
s γ = 1 + 4 u γ , γ { 0.80 , 1.00 , 1.25 }
The exponent γ = 0.80 yields a permissive concave membership function, γ = 1.00 reproduces the delivered linear transformation, and γ = 1.25 yields a conservative convex function. The bounding exponents were selected a priori as reciprocal deviations around the linear case, since 1.25 = 1 0.80 . They provide a moderate, log-symmetric perturbation of membership function curvature while preserving the scale endpoints and the ordering within each layer. These bounds were defined as part of the robustness analysis and should not be interpreted as confidence limits or estimates of measurement error. Applying the three alternatives independently to the four geoscientific layers produced 3 4 = 81 evidence models, after which zonal means and both Pareto frontiers were recomputed.
Transformation robustness was summarized using two outputs. Evidence model frontier frequency F T is the fraction of the 81 scenarios in which an area is nondominated under the baseline group weights, while Jaccard similarity measures retention of frontier identity relative to the linear baseline. Joint robustness frequency F J was then estimated from 10,000 Monte Carlo replicates by pairing one expert bootstrap weight vector with one transformation scenario selected uniformly from the 81 scenarios and recomputing both frontiers. Equal scenario selection gives each of the 3 4 = 81 cells of the factorial stress test the same representation; it does not assign an empirical probability to any transformation model. Expert-only bootstrap membership frequency F E remains a separate output. Because the three membership functions are stress test cases rather than calibrated stochastic models, F T and F J are robustness frequencies and should not be interpreted as empirical probabilities of spatial layer errors. In each joint replicate, the expert bootstrap draw and the transformation scenario were selected independently. This independence is a transparent design assumption because this study provides no empirical basis for associating an expert’s preferences with a particular transformation model. Accordingly, F J is conditional on this independence assumption and would require reinterpretation if such dependence were established.

2.8. Secondary Action Labels and Environmental Retention

For communication after the Pareto analysis, secondary action labels were defined from the sample medians of geoscientific prospectivity, territorial compatibility, and evidence gap. Values equal to a median were assigned to the upper side of that dimension; thus, ‘high’ means greater than or equal to the median, and ‘low’ means strictly below it. Areas high in prospectivity and compatibility were labeled advanced candidates when their gap was low, or they were labelled characterize-first opportunities when it was high. Areas high in prospectivity but low in compatibility were labeled conflict-sensitive. Areas low in prospectivity and low in evidence gap were labeled evidence-rich with lower relative prospectivity; the remainder were frontier research targets. Under this rule, the exact median values for Cerro Bravo and Villamaría–Termales are consistently assigned to the upper prospectivity and upper evidence-gap sides, respectively. These median rules are a descriptive communication layer, not a universal cutoff or a portfolio optimization model.
Environmental retention was measured as 100 times the corrected polygon area divided by its original registered area. The national clean surface was screened at the project’s reference threshold of 3.5. Cells above this threshold were intersected with official geothermal polygons and municipal boundaries to identify how much high-priority area remained outside the existing inventory. All raster calculations used a 300 m cell size (0.09 km2). The territorial-diversification mask was a rasterization of the same 23 official geothermal-area polygons used in the overlap comparison, and the estimated potential was assigned through inventory-linked influence zones. The overlap statistic was therefore treated as an internal consistency check, not independent validation. To expose this dependence, the national screen was repeated after removing potential and territorial diversification; renormalizing the remaining nine weights; and applying thresholds of 3.25, 3.50, and 3.75. The polygon correction defines the spatial support for zonal summaries, while the sensitive-ecosystem layer enters once as a proximity criterion within territorial compatibility. Environmental retention is not multiplied into the composite and is not a Pareto objective. The two stages can nevertheless reinforce environmental screening, so robustness to alternative correction geometries remains a validation priority.

3. Results

3.1. Criterion Importance and Expert-Panel Sensitivity

Estimated electric potential was the most influential criterion, with a group weight of 0.1584. Detailed-study evidence followed at 0.1102, while geothermal gradient (0.1021) and fluid temperature (0.1011) had nearly equal group importance. Sensitive ecosystems received 0.0941. The remaining criteria ranged from 0.0589 for government interest to 0.0804 for aquifer evidence. This ordering indicates that the panel did not frame geothermal prioritization as resource magnitude alone: Evidence maturity and environmental compatibility together accounted for substantial decision weight. Figure 1 summarizes the group BWM weights and their 95% expert-resampling percentile intervals across the 11 criteria.
The expert-resampling percentile intervals reveal where apparent precision should be avoided. Estimated potential ranged from 0.1148 to 0.2073 across the central 95% of resampled panels. Detailed-study evidence had an interval of 0.0754–0.1626 and the largest expert standard deviation (0.1263 after individual-vector normalization). Geothermal gradient also showed appreciable variability (0.0742–0.1435). By contrast, territorial diversification and the Indigenous-territory criterion had narrower intervals. The corresponding group weights, expert-resampling percentile intervals, and expert dispersion values are reported in Table 2. These patterns imply that adding experts with different exploration backgrounds could alter mid-ranked areas even when the leading tier remains stable.

3.2. Composite Ranking Is Not the Same as Geoscientific Prospectivity

Cubarral achieved the highest clean composite score (3.806; 95% expert-resampling percentile interval of 3.676–3.933) and ranked first in every bootstrap replicate. Apiay was second (3.600; 3.478–3.720) and also retained its position in every replicate. Nereidas ranked third with a median bootstrap rank of 3 and a 71.5% expert-resampling frequency of appearing in the top three. Iza ranked fourth overall but entered the top three in 28.5% of replicates. The next tier, comprising San Diego, Paipa, Cerro Machín, Azufral, RUMBA, and Santa Rosa, showed progressively wider rank intervals. RUMBA was especially sensitive, with a 95% rank interval from 5 to 14. Figure 2 shows the rank uncertainty of the ten leading registered areas under expert resampling, and Table 3 summarizes their clean composite scores and diagnostic dimensions.
The dimension-separated results change the interpretation of this ordering. Paipa (3.132) and Santa Rosa (3.135) had the highest geoscientific prospectivity scores among the leading areas, despite ranking sixth and tenth in the composite. Iza also scored highly on prospectivity (3.080). Cubarral’s composite leadership instead reflected a balanced combination of moderate-to-high prospectivity, very high territorial compatibility (4.842), strong evidence maturity, institutional alignment, and full environmental-area retention. Apiay had high compatibility and mature evidence but fell just below the sample median for geoscientific prospectivity. The composite is therefore useful as a descriptive screening reference, but it must not be reported as a geological resource ranking.

3.3. Advancement and Learning Frontiers

Five areas formed the baseline advancement frontier: Cubarral, Nereidas, Iza, Paipa, and Santa Rosa. Cubarral, Nereidas, and Paipa remained advancement-efficient in every expert bootstrap; Iza did so in 99.68% of replicates and Santa Rosa in 84.03%. The result is not a five-site ranking. Each area survives because no competitor simultaneously improves prospectivity and territorial compatibility while reducing its evidence gap. Nereidas, for example, remains efficient because its combination of prospectivity and mature evidence is not replicated elsewhere, despite comparatively low compatibility. Santa Rosa is less stable because modest changes in dimension weights can make it dominated by a nearby alternative. Figure 3 shows the advancement frontier in the prospectivity, compatibility, and evidence-gap space.
Nine areas formed the baseline learning frontier. Cubarral, Iza, San Diego, Paipa, Santa Rosa, Chiles–Cerro Negro, Nevado del Tolima, Doña Juana–Las Ánimas, and Huila were non-dominated when a large evidence gap was treated as an information-acquisition screening signal. San Diego, Chiles–Cerro Negro, and Huila had 100% learning-front membership; Nevado del Tolima had 99.39%, Paipa had 97.77%, Doña Juana–Las Ánimas had 94.63%, and Santa Rosa had 61.31%. Paletará was not on the baseline frontier but entered it in 10.07% of resampled panels. These empirical frequencies distinguish targets with persistent non-dominance under expert-panel resampling from cases for which their learning-front membership depends on expert composition. The corresponding learning frontier is also shown in Figure 3.
The contrast with the weighted ranking is consequential. Apiay occupied the second composite position in every bootstrap replicate, yet it belonged to neither Pareto frontier because other areas offered a more efficient combination of prospectivity, compatibility, and evidence status. Conversely, Huila ranked near the lower end of the composite but remained on the learning frontier because its evidence deficit represents a distinct unresolved exploration question. A stable composite rank and a stable decision frontier therefore answer different questions. Table 4 summarizes the baseline frontier roles and their robustness under expert resampling, geoscientific transformation scenarios, and the joint ensemble.

3.4. Robustness to Geoscientific-Transformation Uncertainty

The linear cellwise scenario reproduced the delivered zonal means to machine precision. The advancement frontier was identical in all 81 evidence models (Jaccard = 1.000): Cubarral, Nereidas, Iza, Paipa, and Santa Rosa each had TF (advancement) = 100%. The learning frontier was also highly stable, with Jaccard similarity from 0.889 to 1.000 (median 0.900) and exact reproduction of the baseline set in 39 scenarios. Eight of the nine baseline learning members persisted in every scenario. Santa Rosa was the only baseline member sensitive to transformation shape, with TF (learning) = 51.9%; Cerro Machín entered in 3.7% of scenarios. The changes therefore concerned a boundary case rather than wholesale replacement of the decision set.
The joint ensemble did not overturn the core conclusion. Joint advancement robustness frequency was 100% for Cubarral and Nereidas, 99.96% for Paipa, 98.84% for Iza, and 83.38% for Santa Rosa. For learning, Cubarral, Iza, San Diego, Chiles–Cerro Negro, and Huila retained 100% joint robustness frequency; Nevado del Tolima, Doña Juana–Las Ánimas, and Paipa remained above 92.9%, while Santa Rosa declined to 58.2%. Paletará and Cerro Machín reached only 15.3% and 15.1%, respectively, identifying them as jointly sensitive boundary cases rather than robust frontier members. Table 4 and Figure 4 thus distinguish preference sensitivity, transformation sensitivity, and joint sensitivity at the decision endpoint.
The missing-data sensitivity separated a stable core from rule-sensitive boundary cases. Conservative score-1 imputation reproduced the five-member advancement frontier exactly (Jaccard = 1.000); its learning frontier had Jaccard = 0.800 because Santa Rosa left and Gabriel López entered. Neutral score-3 imputation retained Cubarral, Nereidas, and Santa Rosa on the advancement frontier but removed Iza and Paipa (Jaccard = 0.600); the learning frontier retained eight of nine baseline members, with only Paipa leaving (Jaccard = 0.889). Thus, the principal advancement and learning signals persisted, but some frontier identities depend on how incomplete geoscientific coverage is incorporated. These nonspatial endpoints are reported in Supplementary File S6. Varying α across 0.30, 0.50, and 0.70 left both frontier identities unchanged, with advancement and learning Jaccard similarities of 1.000 in every specification. Supplementary File S8 releases the area-level evidence gaps and frontier endpoints.

3.5. Secondary Exploration-Action Labels

Three areas were classified as advanced exploration candidates: Cubarral, Paipa, and Santa Rosa. They combine above-median geoscientific prospectivity with above-median territorial compatibility and a below-median evidence gap. The label does not imply equivalent project maturity. Cubarral is a registered area associated with subsurface hydrocarbon information and full spatial retention; Paipa and Santa Rosa require very different technical and environmental programs. The value of the class is that new spending can emphasize prospect testing and project-specific validation rather than basic regional data acquisition.
Four areas were characterize-first opportunities: Iza, San Diego, Azufral, and Villamaría–Termales. Their relative resource and compatibility conditions justify continued attention, but evidence gaps remain above the median. The appropriate next step is targeted information gain: geochemical sampling, updated structural interpretation, magnetotellurics or other geophysical surveys, and carefully staged drilling where warranted. Five areas—Nereidas, Cerro Machín, Chiles–Cerro Negro, Cerro Bravo, and Paletará—were conflict-sensitive resources because their prospectivity was above the median while compatibility was below it. For these areas, technical exploration should be coupled with early environmental and rights-based procedural work; advancing subsurface knowledge without addressing territorial conditions would leave a major implementation risk unresolved.
Four areas were evidence-rich, with lower relative prospectivity: Apiay, RUMBA, Cumbal, and Sotará–Sucubún. These sites should not be promoted as the strongest resource prospects merely because their evidence status is comparatively favorable. For Apiay and RUMBA, this reflects mature subsurface evidence, whereas for Cumbal and Sotará–Sucubún, it primarily reflects high technical data coverage. They may nevertheless provide useful analogue, monitoring, or alternative use opportunities, but their resource concept must be matched to temperature and end-use requirements. Seven areas were frontier research targets. These require low-cost reconnaissance and conceptual model development before they compete for expensive exploration funds. Figure 5 shows the distribution of the 23 registered areas in terms of geoscientific prospectivity and evidence gap, with bubble size proportional to the estimated MWe and color indicating the secondary action label. Table 5 summarizes the secondary exploration action classes and the corresponding decision implications.
The action labels were stable to the evidence-gap mixing assumption. In total, 19 of 23 areas retained the same label under the 30/70, 50/50, and 70/30 combinations of study maturity and technical coverage. Changes were limited to cases near the median gap boundary; no area moved between a conflict-sensitive label and a compatibility-favorable label because territorial compatibility was independent. This supports the labels for communication, while the continuous indicators and frontier robustness frequencies should govern decisions near boundaries. Because the medians are calculated from the current 23-area comparison set, the labels are comparison-dependent and may change when new areas are added even if an existing area’s measurements do not. They should therefore be used as communication descriptors rather than eligibility thresholds; continuous diagnostics, frontier membership, and robustness endpoints remain the decision basis.

3.6. Environmental Retention and Spatial Concentration

Environmental correction substantially reduced the spatial footprint of many recognized geothermal areas. The median retained share was 53.4%, and 11 of the 23 polygons retained less than half of their original area. Cumbal retained 12.9%, Santa Rosa retained 14.7%, Nevado del Tolima retained 16.3%, Sotará–Sucubún retained 25.3%, Paletará retained 30.5%, and Paipa retained 46.9%. In contrast, Cubarral, Apiay, Nereidas, and Iza retained their full mapped areas. The retention metric should not be read as an environmental-impact assessment; it quantifies the spatial consequence of applying the project’s national sensitivity correction and shows where site-scale alternatives are likely to be narrow. Figure 6 shows the relationship between the retained area after environmental correction and the clean composite score, with bubble size proportional to the estimated MWe. The denominator is the original registered geometry, not a reconstructed pre-correction suitability surface; consequently, low retention can reflect both environmental exclusions and an initially broad or overestimated boundary. The metric does not assume that every original cell was equally developable.
The clean national surface contained 109,169 valid screening cells, equivalent to 9825.21 km2. Under the 11-criterion model, the share of high-score cells inside registered polygons was 99.5% at 3.25, 99.7% at 3.50, and 100% at 3.75. At the reference threshold of 3.50, 3685 cells (331.65 km2) were selected, including 12 cells (1.08 km2) outside the polygons in Becerril, Cesar. After removing estimated potential and territorial diversification and renormalizing the remaining nine weights, the corresponding inside-polygon shares were 55.2%, 69.6%, and 85.3%; at 3.50, 996 of 3280 high-score cells (89.64 km2) occurred outside the registered inventory. The threshold-only result is stable, but the reduced-model result confirms that the original 99.7% largely reflects inventory-linked criteria. It is therefore an internal consistency diagnostic rather than independent validation of the official polygons. The Becerril signal remains a baseline data-check and reconnaissance cue, not evidence of a new developable field. Table 6 reports the full comparison, and Figure 7 maps the baseline threshold of 3.50.

4. Discussion

4.1. From Favorable Cells to Exploration Decisions

The central result is interpretive rather than cartographic: a single high geothermal score can arise from different combinations of resource promise, evidence maturity, territorial compatibility, and policy alignment. This distinction is consistent with geothermal play-fairway thinking, which treats evidence and confidence as related but non-identical properties [9]. It is also essential in Colombia, where one deep or hydrocarbon well can sharply improve knowledge without establishing the strongest hydrothermal prospect and where a high-potential volcanic area may face a narrow environmentally compatible footprint.
The contrast between Apiay and Iza illustrates the practical value. Apiay is second in the composite and extremely rank-stable because it combines well-related evidence and strong compatibility; nevertheless, its geoscientific score is comparatively lower, and it is dominated on both decision frontiers. Iza ranks fourth, has higher geoscientific prospectivity and compatibility, and a large evidence gap; it remains efficient for both advancement and learning in essentially every resampled panel. A conventional ranking might allocate the next dollar to Apiay because it is second. The Pareto analysis instead distinguishes whether a candidate remains non-dominated for advancement or for preliminary information-gap screening; it does not identify the optimal expenditure. Apiay may still offer analogue, monitoring, or lower-temperature-use value, but the composite rank alone does not establish priority.
Nereidas provides a second example. It remains in the leading tier and has no evidence gap under the implemented indicator, consistent with its deep exploration history. Yet it falls below the territorial-compatibility median. The appropriate conclusion is neither to discard the resource nor to equate knowledge with readiness. Rather, technical work and territorial-risk management must be designed as a coupled program. This is especially important because national-scale environmental and sociocultural layers are screening inputs; only site-specific engagement can establish legitimate project alternatives.

4.2. Methodological Contribution

This study combines four methodological elements that together support the decision framework. First, it treats semantic auditing as part of reproducible GIS–MCDA. Layer names, transformation direction, and criterion multiplicity were inspected before calculation. This step exposed a consequential ambiguity in the detailed-study criterion: Its title suggested scarcity, whereas its implemented scale measured maturity. Making the direction explicit prevents opposite policy interpretations of the same raster.
Second, the analysis reduces structural double counting. When the expert panel weights one conceptual criterion, administrative or cartographic sublayers should not each receive the full weight unless the elicitation explicitly distinguished them. Averaging municipal and departmental interest, consolidating land representations, and selecting between alternative gradient surfaces preserve the elicited 11-criterion decision structure. This is a simple but transferable audit rule for spatial MCDA.
Third, uncertainty is propagated to decision outputs rather than stopping at weight intervals. The expert bootstrap quantifies whether an area remains non-dominated under changes in panel composition, while the cellwise transformation stress test asks whether the same decision survives bounded changes in how the four dominant geoscientific layers translate evidence into standardized scores. Their joint ensemble makes the interaction observable. The unchanged advancement set and the localized learning-front changes show that preference uncertainty and evidence-model uncertainty need not have the same decision consequences.
Fourth, the paired Pareto formulation operationalizes non-compensation without abandoning the transparent weighted sum as a diagnostic reference. No site is recommended from that score alone, and no cross-dimension exchange rate is imposed. The advancement frontier searches for efficient combinations of promise, compatibility, and low evidence risk; the learning frontier treats unresolved evidence as a reason to investigate rather than as proof of low resource values. Unlike suitability-oriented implementations that terminate in a composite score or favorability class, the paired-frontier framework preserves two non-equivalent decisions: advancing comparatively mature prospects and acquiring information where uncertainty remains consequential.

4.3. Position Within Geothermal-Resource Assessment

Geothermal GIS–MCDA has progressively improved the integration of geological, thermal, geochemical, environmental, and accessibility evidence, from regional favorability mapping to play-fairway and depth-integrated assessments [4,5,6,7,8,9]. Related applications address mine-land direct use, hot-spring end uses, information-driven regional screening, and alternative aggregation operators [10,11,25,26]. These advances improve where and how prospectivity is estimated. The present framework addresses the next decision layer: after screening, which alternatives cannot be improved simultaneously in prospectivity, territorial compatibility, and evidence status?
Geothermal-risk research has separately emphasized multiple subsurface interpretations, decision analysis, and value-of-information (VOI) methods [1]. Recent play-based work monetizes how information obtained from one prospect changes the economics of subsequent prospects through spatial correlation and decision trees [2]. The learning frontier proposed here is deliberately earlier-stage and non-monetary. It identifies registered areas where a detailed VOI analysis may be consequential when comparable survey costs, success probabilities, and cash-flow models are not yet available at the national scale; it is not a substitute for economic VOI.
This positioning is relevant to resource management because scarce public attention must be divided between advancing comparatively mature areas and resolving uncertainty in areas that may otherwise remain systematically undervalued. The two objectives require different evidence directions and different decision gates, even when they draw on the same national database.

4.4. Contribution to Exploration-Decision Theory

The paired-frontier formulation contributes three decision principles. First, the meaning of an evidence gap is made conditional on the objective. In an advancement decision, a large gap is a liability; in a learning decision, it can be an opportunity, but only when the area’s prospectivity and territorial compatibility are not jointly dominated. The learning frontier therefore does not reward ignorance indiscriminately.
Second, Pareto dominance prevents unexamined compensation across prospectivity, compatibility, and evidence status. The weighted composite remains useful for descriptive screening, but frontier membership answers a stricter question: whether another observed area performs at least as well on every decision dimension and better on at least one. Apiay and Huila demonstrate the consequence. Apiay is a perfectly stable composite leader but is dominated on both frontiers; Huila is low in the composite yet persistently efficient for learning.
Third, frontier-membership and robustness frequencies convert a static non-dominated set into design-conditional stability measures under two uncertainty sources. Values near one across expert, transformation, and joint endpoints indicate robust non-dominance under the examined specifications; divergence among the three identifies whether status is preference-sensitive, evidence-model-sensitive, or jointly sensitive. These outputs can identify candidates for subsequent costed appraisal, but they do not allocate funds, establish net information value, or replace project-scale go/no-go analysis.

4.5. Limitations and Validation Priorities

Several limitations bound the claims. The 300 m resolution is appropriate for regional screening but not for well siting. Thermal springs can be spatially offset from their underlying reservoirs, and temperature interpolation within a 5 km envelope can create apparent continuity unsupported by permeability or structural connectivity. The national gradient surface has incomplete coverage and relies largely on hydrocarbon bottom-hole temperatures. Aquifer presence is a coarse proxy for fluid availability and does not establish reservoir transmissivity. Estimated MWe values inherit the assumptions of the national volumetric assessment [3,17,27].
The BWM bootstrap quantifies sensitivity to the composition of the observed expert panel, and the new cellwise stress test addresses bounded curvature in the four dominant standardized transformations. Neither analysis quantifies measurement error, spatial misregistration, interpolation uncertainty, or uncertainty in the raw MWe, temperature, gradient, and thermal-spring observations; the transformation exponents are stress-test bounds rather than estimated probability distributions. Pareto membership also depends on the selected dimensions and on the observed 23-area comparison set, and it can change when new registrations or evidence are included. The learning frontier is a preliminary information-gap screen and does not estimate value of information, survey cost-effectiveness, or expected decision value. Government-interest coding measures explicit policy attention, not budget execution, permitting capacity, or long-term commitment. Distance from Indigenous territories must not be interpreted as acceptance or opposition; it is only a national-scale indicator of procedural complexity, and project decisions require rights-based participation and consultation. The joint ensemble further assumes independent pairing of preference and transformation draws. Neither the expert bootstrap nor the original elicitation distinguishes criterion-specific expertise or confidence. The environmental-retention denominator is an administratively registered polygon and may inherit initially broad boundaries. Secondary median labels are comparison-set dependent. Finally, the learning frontier does not include spatial contiguity, infrastructure, survey logistics, or cross-area information spillovers; isolated members should not be interpreted as immediately executable survey programs. All robustness frequencies are conditional on the observed 23-area comparison set, selected objectives, α range, transformation bounds, equal factorial scenario design, and independence assumption. The protected geodatabase prevents independent spatial rerunning, and some institutional source-layer snapshot dates were not retained in the analytical export; the implemented rules are auditable in S5 but cannot substitute for validation against the restricted source GIS.
The next validation stage should prioritize four additions: (i) independent geochemistry and geothermometry for characterize-first areas; (ii) structural and permeability evidence, including fault characterization and geophysics; (iii) empirical error models and spatial ensembles for the raw MWe, temperature, gradient, and thermal-spring inputs; and (iv) participatory validation of territorial compatibility in Paipa, Villamaría, and other affected areas. A value-of-information analysis could then estimate which survey most improves the probability of an exploration decision at the lowest cost. A portfolio implementation should additionally cluster compatible surveys, represent access and infrastructure constraints, elicit criterion-specific expert confidence, and test dependence structures between preference and evidence-model uncertainty.

4.6. Transferability, Resource Outlook, and Implications

The framework is transferable as a decision architecture rather than as a fixed Colombian scoring model. Its required elements are a common set of portfolio units, locally meaningful prospectivity and compatibility criteria, an explicit representation of evidence maturity and spatial coverage, and a defensible comparison set. In volcanic, sedimentary, enhanced-geothermal, or direct-use settings, the evidence layers, distance functions, institutional indicators, rights context, and decision thresholds must be re-elicited. The Colombian BWM weights, 5 km envelope, median labels, and resource cutoffs should not be exported without local validation. What transfers is the sequence: audit semantics and duplicated representations, keep evidence status separate from resource favorability, compute advancement and learning frontiers, and stress-test both preference and evidence models [12,13,14,15,16].
This architecture is particularly relevant where exploration data are spatially uneven. A conventional favorability map can systematically demote poorly studied areas; the paired frontiers instead expose whether an area is efficient under a low-evidence-risk objective or under an information-acquisition objective. Implementation elsewhere should report the local scoring rules; distinguish missing observations from unfavorable evidence; and add cost, accessibility, infrastructure, contiguity, permitting, and community-governance constraints before converting a learning frontier into a field program.
For Colombia, the approximately 1170.2 MWe reported by the national volumetric assessment remains contextual inventory potential, not an updated estimate of recoverable reserves, generation capacity, or bankable projects. The present analysis does not revise that resource magnitude. Its outlook is procedural: robust advancement-front members are candidates for project-scale conceptual modeling and gated validation, whereas robust learning-front members identify areas where targeted data acquisition merits subsequent costed evaluation. The reduced-criterion national sensitivity also shows that high-score concentration inside official polygons is partly constructed by inventory-linked inputs; it cannot validate those polygons independently. New outside-polygon signals should therefore be treated as reconnaissance hypotheses, while registered areas still require reservoir confirmation, environmental alternatives, rights-based engagement, and economic appraisal before development decisions.

5. Conclusions

  • This study converts a national geothermal GIS–BWM screening model into an evidence-aware, non-compensatory decision framework by reporting prospectivity, compatibility, institutional alignment, and evidence status separately and distinguishing efficient advancement from efficient learning.
  • Cubarral, Nereidas, Iza, Paipa, and Santa Rosa formed the baseline advancement frontier, which was unchanged across all 81 transformation scenarios. Eight of nine learning-front members persisted across every transformation scenario, while Santa Rosa was the principal transformation-sensitive case. Apiay’s invariant second composite rank but absence from both frontiers confirms that rank stability is not decision efficiency.
  • The missing-data tests retained a stable core but identified boundary sensitivity: conservative score-1 imputation reproduced the advancement set, whereas neutral score-3 imputation retained three of its five members; learning-front Jaccard similarities were 0.800 and 0.889. The 99.7% baseline overlap with official polygons is robust to threshold changes but falls to 69.6% at the 3.50 threshold after removing potential and diversification, so it is an internal consistency check rather than independent validation.
  • For resource agencies, the operational implication is to maintain separate analytical queues and decision gates for advancement and learning, while allocation still requires cost, feasibility, and expected-value analysis. The architecture can be transferred to other regions only after local redefinition of evidence layers, transformations, institutional and rights indicators, and feasibility constraints; it structures the next comparison but does not estimate recoverable reserves or authorize development.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/resources15090119/s1. S1_geothermal_area_decision_results.csv, containing area-level scores, uncertainty intervals, diagnostics, evidence indicators, action labels, environmental retention, Pareto membership, and robustness metrics; S2_criterion_weights_summary.csv, containing aggregated BWM weights, expert dispersion, and expert-resampling percentile intervals; S3_transformation_frontier_membership.csv, containing membership across the 81 geoscientific transformation scenarios and Jaccard metrics; S4_decision_analysis.py, containing the public nonspatial verification workflow; S5_criterion_scoring_rules.csv, containing original variables, sources and years, spatial preprocessing, numerical 1–5 rules, categorical assignments, source-consolidation rules, two-source gradient precedence, and missing-value treatment for all 11 criteria; S6_missing_data_rule_frontiers.csv, containing area-level frontier endpoints under conservative and neutral missing-cell imputations; S7_national_surface_sensitivity.csv, containing aggregated threshold and reduced-criterion overlap results; S8_alpha_frontier_sensitivity.csv, containing area-level evidence gaps and both frontier identities under α = 0.30, 0.50, and 0.70; and README.md, containing the package description, data dictionaries, methodological constants, verification instructions, and access limitations.

Author Contributions

Conceptualization, C.R.-A., V.O.-O., and C.D.R.; methodology, C.R.-A. and V.O.-O.; formal analysis, C.R.-A.; software, C.R.-A.; validation, C.R.-A., V.O.-O., and C.D.R.; investigation, C.R.-A., V.O.-O., and C.D.R.; resources, V.O.-O. and C.D.R.; data curation, C.R.-A. and V.O.-O.; writing—original draft preparation, C.R.-A.; writing—review and editing, C.R.-A., V.O.-O., and C.D.R.; visualization, C.R.-A.; supervision, V.O.-O. and C.D.R.; project administration, V.O.-O. and C.D.R.; funding acquisition, C.D.R. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Agencia Nacional de Hidrocarburos (ANH) through Contract No. 618 of 2025, implemented by the Universidad del Magdalena within the GENTE project, Gobernanza ENergética y TErritorio. The project addresses the diagnosis, analysis, and development of methodological guidelines to support energy transition planning and territorial decision-making. The article processing charge was also covered by the Agencia Nacional de Hidrocarburos (ANH).

Institutional Review Board Statement

This study involved a non-interventional technical expert-elicitation questionnaire using the Best–Worst Method (BWM). Based on Article 2 of Rectoral Resolution No. 427 of 6 July 2018 of Universidad del Magdalena, which defines the scope of the university’s Research Ethics Committee, the authors considered that prior ethics committee approval was not required. The questionnaire was limited to obtaining professional judgments on the relative importance of geothermal planning criteria and did not involve interventions affecting human or non-human health or life, the environment, or biodiversity.

Informed Consent Statement

Informed consent was obtained from all experts participating in this study. Before completing the questionnaire, participants authorized the processing of their personal data and were informed that the information collected would be used solely for technical and academic purposes and that the results would be reported in aggregate form without associating individual responses with specific participants.

Data Availability Statement

The nonspatial aggregated results and verification code supporting the principal findings of this study are provided in the Supplementary Materials. The original project geodatabase and spatially explicit source or derived data, including raster and vector layers, geometries, coordinates, and cell-level outputs, are not publicly available because they contain institutional and third-party data subject to access and redistribution restrictions. Access to these materials is subject to authorization by the Agencia Nacional de Hidrocarburos (ANH) and Universidad del Magdalena. Individual expert responses are not shared; only aggregated statistics are reported, consistent with participant consent.

Acknowledgments

The authors gratefully acknowledge the geothermal experts who participated in the BWM elicitation process, as well as the technical teams of the Agencia Nacional de Hidrocarburos (ANH) and Universidad del Magdalena for their contributions to the compilation, standardization, and documentation of the GENTE geospatial database.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Witter, J.B.; Trainor-Guitton, W.J.; Siler, D.L. Uncertainty and risk evaluation during the exploration stage of geothermal development: A review. Geothermics 2019, 78, 233–242. [Google Scholar] [CrossRef] [Scilit]
  2. van Unen, M.; Brunner, L.; Veldkamp, H.; Keijzer, R.; van Wees, J.D. De-risking of geothermal prospect portfolios based on Value of Information and a play-based exploration approach: A case study for Slochteren reservoir development in the Netherlands. Glob. Planet. Change 2026, 256, 105130. [Google Scholar] [CrossRef] [Scilit]
  3. Muffler, L.J.P.; Cataldi, R. Methods for regional assessment of geothermal resources. Geothermics 1978, 7, 53–89. [Google Scholar] [CrossRef] [Scilit]
  4. do Amaral, M.L.; Caldeira, M.C.O.; de Figueiredo, J.J.S.; da Silveira, J.R.B. Integration of geophysical data and multicriteria decision analysis for geothermal assessment at Utah FORGE. Geothermics 2026, 136, 103590. [Google Scholar] [CrossRef] [Scilit]
  5. Jara-Alvear, J.; De Wilde, T.; Asimbaya, D.; Urquizo, M.; Ibarra, D.; Graw, V.; Guzmán, P. Geothermal resource exploration in South America using an innovative GIS-based approach: A case study in Ecuador. J. S. Am. Earth Sci. 2023, 122, 104156. [Google Scholar] [CrossRef] [Scilit]
  6. Noorollahi, Y.; Itoi, R.; Fujii, H.; Tanaka, T. GIS model for geothermal resource exploration in Akita and Iwate prefectures, northern Japan. Comput. Geosci. 2007, 33, 1008–1021. [Google Scholar] [CrossRef] [Scilit]
  7. Yalcin, M.; Kilic Gul, F. A GIS-based multi-criteria decision analysis approach for exploring geothermal resources: Akarcay basin (Afyonkarahisar). Geothermics 2017, 67, 18–28. [Google Scholar] [CrossRef] [Scilit]
  8. Meng, F.; Liang, X.; Xiao, C.; Wang, G. Geothermal resource potential assessment utilizing GIS-based multi-criteria decision analysis method. Geothermics 2021, 89, 101969. [Google Scholar] [CrossRef] [Scilit]
  9. DeAngelo, J.; Shervais, J.W.; Glen, J.M.; Nielson, D.; Garg, S.; Dobson, P.F.; Gasperikova, E.; Sonnenthal, E.; Liberty, L.M.; Siler, D.L.; et al. Geothermal Play Fairway Analysis, Part 2: GIS methodology. Geothermics 2024, 117, 102882. [Google Scholar] [CrossRef] [Scilit]
  10. Gasperikova, E.; Ulrich, C.; Omitaomu, O.A.; Dobson, P.; Zhang, Y. Multicriteria screening evaluation of geothermal resources on mine lands for direct use heating. Geotherm. Energy 2024, 12, 11. [Google Scholar] [CrossRef] [Scilit]
  11. Ng’ethe, J.; Jalilinasrabady, S. GIS-based AHP model for selecting the best direct-use scenarios for medium- to low-enthalpy geothermal resources with hot springs in central and western Kenya. Geothermics 2024, 122, 103069. [Google Scholar] [CrossRef] [Scilit]
  12. Shervais, J.W.; DeAngelo, J.; Glen, J.M.; Nielson, D.L.; Garg, S.; Dobson, P.F.; Gasperikova, E.; Sonnenthal, E.; Liberty, L.M.; Newell, D.L.; et al. Geothermal play fairway analysis, part 1: Example from the Snake River Plain, Idaho. Geothermics 2024, 117, 102865. [Google Scholar] [CrossRef] [Scilit]
  13. Siler, D.L.; Zhang, Y.; Spycher, N.F.; Dobson, P.F.; McClain, J.S.; Gasperikova, E.; Zierenberg, R.A.; Schiffman, P.; Ferguson, C.; Fowler, A.; et al. Play-fairway analysis for geothermal resources and exploration risk in the Modoc Plateau region. Geothermics 2017, 69, 15–33. [Google Scholar] [CrossRef] [Scilit]
  14. Ito, G.; Frazer, N.; Lautze, N.; Thomas, D.; Hinz, N.; Waller, D.; Whittier, R. Play fairway analysis of geothermal resources across the state of Hawaii: 2. Resource probability mapping. Geothermics 2017, 70, 393–405. [Google Scholar] [CrossRef] [Scilit]
  15. Lautze, N.; Thomas, D.; Waller, D.; Frazer, N.; Hinz, N.; Apuzen-Ito, G. Play fairway analysis of geothermal resources across the state of Hawaii: 3. Use of development viability criterion to prioritize future exploration targets. Geothermics 2017, 70, 406–413. [Google Scholar] [CrossRef] [Scilit]
  16. Athens, N.D.; Caers, J.K. A Monte Carlo-based framework for assessing the value of information and development risk in geothermal exploration. Appl. Energy 2019, 256, 113932. [Google Scholar] [CrossRef] [Scilit]
  17. Alfaro, C.; Rueda-Gutiérrez, J.B.; Casallas, Y.; Rodríguez, G.; Malo, J. Approach to the geothermal potential of Colombia. Geothermics 2021, 96, 102169. [Google Scholar] [CrossRef] [Scilit]
  18. Alfaro, C.; Alvarado, I.; Quintero, W. Mapa Preliminar de Gradientes Geotérmicos (Método BHT); scale 1:1,500,000; INGEOMINAS-ANH: Bogotá, Colombia, 2008.
  19. Instituto de Hidrología, Meteorología y Estudios Ambientales (IDEAM). Estudio Nacional Del Agua 2022; IDEAM: Bogotá, Colombia, 2023.
  20. Rezaei, J. Best-worst multi-criteria decision-making method. Omega 2015, 53, 49–57. [Google Scholar] [CrossRef] [Scilit]
  21. Liang, F.; Brunelli, M.; Rezaei, J. Consistency issues in the best worst method: Measurements and thresholds. Omega 2020, 96, 102175. [Google Scholar] [CrossRef] [Scilit]
  22. Wu, Q.; Liu, X.; Zhou, L.; Qin, J.; Rezaei, J. An analytical framework for the best-worst method. Omega 2024, 123, 102974. [Google Scholar] [CrossRef] [Scilit]
  23. Mohammadi, M.; Rezaei, J. Bayesian best-worst method: A probabilistic group decision-making model. Omega 2020, 96, 102075. [Google Scholar] [CrossRef] [Scilit]
  24. Tervonen, T.; Lahdelma, R. Implementing stochastic multicriteria acceptability analysis. Eur. J. Oper. Res. 2007, 178, 500–513. [Google Scholar] [CrossRef] [Scilit]
  25. Wang, G.; Xiao, C.; Liang, X.; Deng, Q. Assessment of geothermal resource potential based on GIS information-driven model: A case study of the Songyuan, China. Geothermics 2026, 136, 103587. [Google Scholar] [CrossRef] [Scilit]
  26. Kiavarz, M.; Jelokhani-Niaraki, M. Geothermal prospectivity mapping using GIS-based Ordered Weighted Averaging approach: A case study in Japan’s Akita and Iwate provinces. Geothermics 2017, 70, 295–304. [Google Scholar] [CrossRef] [Scilit]
  27. Franco, A.; Vaccaro, M. Sustainable sizing of geothermal power plants: Appropriate potential assessment methods. Sustainability 2020, 12, 3844. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Group BWM weights and 95% expert-resampling percentile intervals (10,000 resamples).
Figure 1. Group BWM weights and 95% expert-resampling percentile intervals (10,000 resamples).
Resources 15 00119 g001
Figure 2. Rank uncertainty for the ten leading registered areas. Intervals reflect resampling of the 15-expert panel.
Figure 2. Rank uncertainty for the ten leading registered areas. Intervals reflect resampling of the 15-expert panel.
Resources 15 00119 g002
Figure 3. Non-compensatory geothermal exploration frontiers. Filled circles represent the 23 registered geothermal areas and are colored according to evidence gap. Teal circles identify advancement-frontier members, while orange squares identify learning-frontier members; areas without either outline belong to neither baseline frontier. Labels are shown for frontier members and selected reference areas.
Figure 3. Non-compensatory geothermal exploration frontiers. Filled circles represent the 23 registered geothermal areas and are colored according to evidence gap. Teal circles identify advancement-frontier members, while orange squares identify learning-frontier members; areas without either outline belong to neither baseline frontier. Labels are shown for frontier members and selected reference areas.
Resources 15 00119 g003
Figure 4. Frontier robustness under preference and geoscientific-transformation uncertainty. Expert-bootstrap columns show EF from 10,000 resampled panels; transformation columns show TF across 81 equally weighted cellwise membership-function scenarios; joint columns show JRF from paired expert-transformation draws. Values are percentages; TF and JRF are design-conditional robustness frequencies, not empirical probabilities of transformation-model occurrence or population-level inference.
Figure 4. Frontier robustness under preference and geoscientific-transformation uncertainty. Expert-bootstrap columns show EF from 10,000 resampled panels; transformation columns show TF across 81 equally weighted cellwise membership-function scenarios; joint columns show JRF from paired expert-transformation draws. Values are percentages; TF and JRF are design-conditional robustness frequencies, not empirical probabilities of transformation-model occurrence or population-level inference.
Resources 15 00119 g004
Figure 5. Geoscientific prospectivity and evidence gap for the 23 registered areas. Bubble size represents estimated MWe; color represents the secondary action label. The dashed vertical and horizontal lines indicate the sample medians of geoscientific prospectivity (2.730) and evidence gap (54.1%), respectively, used to define the secondary action classes.
Figure 5. Geoscientific prospectivity and evidence gap for the 23 registered areas. Bubble size represents estimated MWe; color represents the secondary action label. The dashed vertical and horizontal lines indicate the sample medians of geoscientific prospectivity (2.730) and evidence gap (54.1%), respectively, used to define the secondary action classes.
Resources 15 00119 g005
Figure 6. Area retained after national environmental correction versus the clean composite score. Bubble size represents estimated MWe. The dashed vertical line indicates the median environmental retention across the 23 registered areas (53.4%).
Figure 6. Area retained after national environmental correction versus the clean composite score. Bubble size represents estimated MWe. The dashed vertical line indicates the median environmental retention across the 23 registered areas (53.4%).
Resources 15 00119 g006
Figure 7. National distribution of clean cells above 3.5 and registered-area frontier roles, with an Andean inset. Symbols identify advancement-only, learning-only, and dual-frontier areas. The isolated Becerril signal (1.08 km2) is shown explicitly. The displayed baseline overlap is an internal consistency view and not independent validation of the inventory.
Figure 7. National distribution of clean cells above 3.5 and registered-area frontier roles, with an Andean inset. Symbols identify advancement-only, learning-only, and dual-frontier areas. The isolated Becerril signal (1.08 km2) is shown explicitly. The displayed baseline overlap is an internal consistency view and not independent validation of the inventory.
Resources 15 00119 g007
Table 1. Criterion structure used in the clean reanalysis.
Table 1. Criterion structure used in the clean reanalysis.
DimensionCriterionOperational RepresentationDirectionImplemented 1–5 Scoring Rule
GeoscientificEstimated power potentialVolumetric MWe estimate assigned to geothermal-area influence zonesHigher is better5: >100 MWe; 4: >5–100 MWe; 3: ≤5 MWe; 2 unused; 1 outside mapped potential/influence or when missing cellwise.
GeoscientificThermal springsDistance to mapped thermal springs; 5 km influenceCloser is better5 at ≤0.1 km; linear decrease for 0.1–5 km; 1 at ≥5 km. score = 5 − 4(d − 0.1)/4.9 inside the interval.
GeoscientificFluid temperatureKriged surface-temperature observations within the evidence envelopeHigher is better1 at ≤31 °C; linear increase for 31–75 °C; 5 at ≥75 °C. score = 1 + 4(T − 31)/44 inside the interval.
GeoscientificGeothermal gradientNational BHT gradient plus area-specific gradient evidenceHigher is better1 at ≤17 °C/km; linear increase for 17–44 °C/km; 5 at ≥44 °C/km. Use area-specific standardized gradient where valid; otherwise, use national BHT gradient; missing only if both are unavailable.
GeoscientificAquifer evidencePresence of mapped aquifer systemsPresence is better5 where mapped aquifer evidence is present; 1 where absent.
Territorial compatibilitySensitive ecosystemsDistance/exclusion relative to protected areas, Ramsar wetlands, and páramosFarther is better1 at ≤0.5 km; linear increase for 0.5–5 km; 5 at ≥5 km, score = 1 + 4(d − 0.5)/4.5 inside the interval.
Territorial compatibilityLand compatibilityConsolidated land-use and mining-use compatibilityCompatible/intervened uses are better5 livestock, extensive/intensive agriculture, or granted mining/extraction; 4 semi-intensive agriculture; 3 limited/difficult agricultural use; 2 unused/unassigned; 1 urban, military, water, protection, conservation, or restoration. Average valid land and mining representations.
Territorial compatibilityIndigenous territoriesDistance to legally recognized Indigenous territoriesFarther is less procedurally complex; not a social-acceptance proxy1 at ≤0.5 km; linear increase for 0.5–5 km; 5 at ≥5 km. This is not an acceptance score.
Institutional alignmentGovernment interestMean of departmental and municipal plan signalsExplicit geothermal/FNCER support is betterEach plan: 5 for an explicit geothermal/non-excluding FNCER signal, otherwise 1; mean of departmental and municipal values (possible 1, 3, or 5).
Institutional alignmentTerritorial diversificationLocation inside an identified/evaluated geothermal areaInside is better5 inside the identified/evaluated geothermal-area polygons; 1 outside.
Evidence maturityDetailed-study evidencePresence of geothermal exploration wells or subsurface hydrocarbon-well evidenceMore direct subsurface evidence is better5 where geothermal exploratory or subsurface hydrocarbon-well evidence is present; 1 otherwise.
Table 2. Group weights, 95% expert-resampling percentile intervals, and dispersion among experts.
Table 2. Group weights, 95% expert-resampling percentile intervals, and dispersion among experts.
CriterionGroup WeightExpert-Resampling 95% IntervalExpert SD
Estimated power potential0.15840.1148–0.20730.0864
Thermal springs0.07180.0564–0.08690.0264
Fluid temperature0.10110.0802–0.12830.0600
Geothermal gradient0.10210.0742–0.14350.1003
Aquifer evidence0.08040.0641–0.09810.0348
Sensitive ecosystems0.09410.0751–0.12350.0784
Land-use compatibility0.06900.0506–0.09010.0348
Indigenous territories0.08030.0682–0.09180.0218
Government interest0.05890.0419–0.08000.0420
Territorial diversification0.07370.0615–0.08420.0174
Detailed-study evidence0.11020.0754–0.16260.1263
Table 3. Ten leading registered areas under the clean composite and their diagnostic dimensions.
Table 3. Ten leading registered areas under the clean composite and their diagnostic dimensions.
RankAreaScoreMedian Rank (Resampling 95% Interval)Expert F (Top 3)ProspectivityCompatibilityEvidence Gap
1Cubarral3.8061 (1–1)100.0%2.8434.8429.8%
2Apiay3.6002 (2–2)100.0%2.7234.4719.8%
3Nereidas3.3023 (3–4)71.5%2.8452.5760.0%
4Iza3.2374 (3–4)28.5%3.0804.00354.1%
5San Diego3.0985 (5–7)0.0%2.9133.87854.8%
6Paipa3.0886 (5–7)0.0%3.1323.35151.2%
7del Volcán Cerro Machín2.9488 (7–12)0.0%2.9632.68251.9%
8del Volcán Azufral2.9428 (7–10)0.0%2.7753.50154.7%
9RUMBA2.9399 (5–14)0.0%2.4192.5699.8%
10del Volcán de Santa Rosa2.9259 (7–12)0.0%3.1352.81251.0%
Table 4. Baseline frontier roles and robustness under expert resampling, geoscientific-transformation scenarios, and the joint ensemble.
Table 4. Baseline frontier roles and robustness under expert resampling, geoscientific-transformation scenarios, and the joint ensemble.
AreaBaseline RoleEF (Adv.)TF (Adv.)JRF (Adv.)EF (Learn.)TF (Learn.)JRF (Learn.)
CubarralBoth100.0%100.0%100.0%100.0%100.0%100.0%
IzaBoth99.7%100.0%98.8%100.0%100.0%100.0%
PaipaBoth100.0%100.0%100.0%97.8%100.0%92.9%
Santa RosaBoth84.0%100.0%83.4%61.3%51.9%58.2%
NereidasAdvancement100.0%100.0%100.0%0.0%0.0%0.0%
San DiegoLearning1.8%0.0%6.2%100.0%100.0%100.0%
Chiles–Cerro NegroLearning0.0%0.0%0.0%100.0%100.0%100.0%
Nevado del TolimaLearning0.0%0.0%0.0%99.4%100.0%99.4%
Doña Juana–Las ÁnimasLearning0.0%0.0%0.0%94.6%100.0%94.6%
HuilaLearning0.0%0.0%0.0%100.0%100.0%100.0%
Cerro MachínNeither0.0%0.0%0.0%9.3%3.7%15.1%
PaletaráNeither0.0%0.0%0.0%10.1%0.0%15.3%
EF: Empirical frequency under 10,000 expert-panel bootstrap resamples; TF: robustness frequency across 81 equally weighted transformation stress tests; JRF: joint robustness frequency under paired expert and transformation draws. TF and JRF are design-conditional frequencies, not calibrated probabilities of transformation-model occurrence or population-level inference.
Table 5. Secondary exploration-action classes and decision implications.
Table 5. Secondary exploration-action classes and decision implications.
ClassNo. of AreasAreasRecommended Next Action
Frontier research target7Nevado del Tolima; Doña Juana–Las Ánimas; Hacienda Granates; Sibundoy; Huila; Gabriel López; Galeras–MorazurcoUse low-cost reconnaissance, data rescue, geochemistry, and conceptual-model building before ranking for drilling.
Conflict-sensitive resource5Nereidas; Cerro Machín; Chiles–Cerro Negro; Cerro Bravo; PaletaráCouple technical work with environmental alternatives, consultation planning, and territorial-risk mitigation.
Evidence-rich, lower prospectivity4Apiay; RUMBA; Cumbal; Sotará–SucubúnReassess resource concept and end use; exploit analogue or monitoring value; avoid overinterpreting data availability.
Characterize-first opportunity4Iza; San Diego; Azufral; Villamaría–TermalesPrioritize information-gain surveys before major capital commitment; close the dominant evidence gaps.
Advanced exploration candidate3Cubarral; Paipa; Santa RosaUpdate conceptual model; validate project-scale constraints; design targeted geophysics and drilling gates.
Table 6. Sensitivity of high-score concentration to the screening threshold and inventory-linked criteria.
Table 6. Sensitivity of high-score concentration to the screening threshold and inventory-linked criteria.
ModelThresholdHigh Cells (km2)Inside Registered Polygons, Cells (km2)Inside (%)
All 11 criteria3.257064 (635.76)7027 (632.43)99.5
All 11 criteria3.503685 (331.65)3673 (330.57)99.7
All 11 criteria3.751458 (131.22)1458 (131.22)100.0
Without potential and diversification3.258052 (724.68)4444 (399.96)55.2
Without potential and diversification3.503280 (295.20)2284 (205.56)69.6
Without potential and diversification3.75948 (85.32)809 (72.81)85.3
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Robles-Algarín, C.; Olivero-Ortiz, V.; Rosas, C.D. Beyond Composite Suitability: Evidence-Aware Pareto Frontiers for Geothermal Exploration Decisions in Colombia. Resources 2026, 15, 119. https://doi.org/10.3390/resources15090119

AMA Style

Robles-Algarín C, Olivero-Ortiz V, Rosas CD. Beyond Composite Suitability: Evidence-Aware Pareto Frontiers for Geothermal Exploration Decisions in Colombia. Resources. 2026; 15(9):119. https://doi.org/10.3390/resources15090119

Chicago/Turabian Style

Robles-Algarín, Carlos, Víctor Olivero-Ortiz, and Carolina Diosa Rosas. 2026. "Beyond Composite Suitability: Evidence-Aware Pareto Frontiers for Geothermal Exploration Decisions in Colombia" Resources 15, no. 9: 119. https://doi.org/10.3390/resources15090119

APA Style

Robles-Algarín, C., Olivero-Ortiz, V., & Rosas, C. D. (2026). Beyond Composite Suitability: Evidence-Aware Pareto Frontiers for Geothermal Exploration Decisions in Colombia. Resources, 15(9), 119. https://doi.org/10.3390/resources15090119

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop