1. Introduction
Digital Elevation Models (DEMs) are elevation datasets derived from a variety of acquisition platforms, sensors, and processing techniques and are widely used in geoscientific applications. While their influence on results has been extensively studied in disciplines such as hydrology [
1,
2,
3], in active tectonics, DEM selection is often limited to resolution considerations, overlooking key aspects such as accuracy and geometric consistency. Not aligning the DEM selection with the study’s requirements can mask or distort results, compromising their reliability [
3,
4,
5,
6,
7,
8]. Relief analysis in active tectonics requires high-precision DEMs to identify subtle morphological variations associated with tectonic deformation [
9,
10]. However, very high-resolution datasets are not always advantageous for regional-scale analyses, where increased sensitivity to local irregularities and topographic noise may hinder the identification of broader geomorphic patterns [
11].
Reliable characterization of terrain morphology in DEMs is critical for ensuring accurate surface analysis. Nevertheless, high-resolution DEMs stay out of reach in many parts of the world. Since the release of the first global DEM, SRTM, in 2005, global models have improved in resolution, accuracy, and correction techniques [
3,
6,
12,
13]. Despite this progress, freely available DEMs with near-global or global coverage are generally limited to spatial resolutions of approximately 30 m, while higher-resolution open access datasets are typically restricted to local or regional scales. This limitation constrains precision and requires researchers to account for error margins in their analyses, underscoring the need to select a freely accessible global DEM whose accuracy and geometric coherence are suitable for ensuring robust results across different applications. Errors in DEMs are not uniform and depend on factors such as topography and data acquisition conditions [
1,
2,
14,
15], which affect their performance in different geographic contexts. Several studies [
15,
16,
17] agree that vertical and horizontal errors are more prevalent in mountainous and/or densely vegetated areas. In particular, ASTER and SRTM tend to show greater inaccuracies in steep terrains [
3,
12,
18], while ALOS has demonstrated better performance in such environments [
12,
15,
16,
17]. This underscores that the significance of DEM errors depends not only on their magnitude but also on context and application, consistent with the fitness-for-use framework [
13,
14]. This is particularly important in coastal geomorphology where precise elevation is critical; small elevation errors can significantly impact the delineation of low-lying coastal zones and the land–sea boundary [
1,
19]. Furthermore, many studies have reported that most global DEMs tend to overestimate terrain elevation [
6,
15], adding another layer of complexity to their use across varied landscapes.
Numerous studies have evaluated open access global DEMs through comparisons with higher-accuracy reference datasets, such as LiDAR-derived models [
15,
20,
21,
22,
23], GPS measurements [
24,
25,
26,
27], or satellite altimetry [
28,
29,
30,
31,
32].
Traditionally, DEM assessments have focused on the quantification of vertical errors and, less frequently, horizontal errors [
21,
26] using global statistical metrics such as RMSE, mean error, median error, or standard deviation. Although these metrics provide a useful evaluation of overall DEM performance, they have important limitations: they assume spatially homogeneous error distributions and do not provide information on the impact of these errors on derived products of error such as slope or curvature [
13,
14].
This limitation is particularly relevant for geomorphological applications, where the usefulness of a DEM depends not only on the absolute accuracy of elevation values but also on the preservation of spatial relationships and terrain morphology [
33,
34]. However, a literature review conducted by Mesa-Mingorance and Ariza López [
35] in 2020 revealed that the majority of studies still focus primarily on global accuracy indicators, while only a limited number assess the spatial variability of errors, and no standardized assessment framework has yet been established. To partially overcome these limitations, recent studies [
12] have shown that comparing terrain derivatives—for example, slope or aspect—is a more comprehensive approach than simply comparing elevation values, as it allows for the analysis of bias and error propagation [
5]. Likewise, some studies have begun to employ stratified analyses based on terrain characteristics, including slope classes [
15,
21,
29,
30,
31,
36,
37], land cover [
15,
20,
22,
26,
28,
29], and landform type [
34,
36,
38].
However, this analysis is often conducted only descriptively [
13] and does not allow evaluation of the extent to which small or large errors may alter the spatial continuity of the landscape and affect the identification and interpretation of landforms, as they do not incorporate fundamental terrain elements that help characterize relief and emphasize its expression [
35]. Consequently, several recent studies have highlighted the need to incorporate methodologies capable of evaluating local topographic coherence and the spatial organization of errors, rather than relying exclusively on point-based statistical metrics [
18,
21,
27,
35].
As a result, there remains a need for methodologies capable of jointly analyzing error magnitude, spatial error distribution, and their implications for the identification and interpretation of landforms.
To address these limitations, this study evaluates the vertical accuracy and geomorphic consistency of freely available global Digital Elevation Models (DEMs), with a particular focus on their suitability for identifying uplifted coastal landforms. High-resolution LiDAR data from El Salvador is used as a reference to assess the limitations of these global DEMs in terms of overall relative vertical accuracy. In addition, the analysis also examines the geometric consistency of derived terrain attributes such as slope. Given the specific interest in detecting uplifted marine terraces to study the active tectonics in subduction-zones coastal settings, a detailed evaluation was designed for the coastal zone (up to 2 km inland), incorporating a density-weighted assessment of spatial error distributions across the terrain. Finally, the uniformity in accuracy and terrain representation capability of the selected DEM is assessed by comparing it with documented uplifted beaches outside the study area.
Beyond the specific interest in evaluating the performance of global DEMs within a particular and geomorphologically heterogeneous region, this study contributes a transferable framework for DEM assessment that integrates three dimensions that are commonly evaluated separately in previous studies: statistical accuracy, spatial error structure and geomorphic coherence. By explicitly linking error patterns to their implications for landform detection, the study reframes DEM evaluation from a purely metric-based comparison toward a process-oriented, application-driven perspective. This approach provides a structured basis for assessing fitness-for-use in geomorphological and tectonic research.
2. Materials and Methods
The methodology was structured in three stages: (1) study area, selection and preprocessing of the DEMs, (2) quantitative evaluation of the global error and analysis of the overall distribution of elevations and slopes, and (3) analysis of error distribution in coastal areas and calculation of the SERI. A high-resolution LiDAR-based DEM corresponding to the coastal area of El Salvador was used as a reference.
2.1. Study Area, Selected DEMs and Preprocessing
The area selected for the study (
Figure 1) covers approximately 31,000 km
2 within the territory of El Salvador, between coordinates 87.7°W–88.4°W in longitude and 13.15°N–13.4°N in latitude. This region includes a wide variety of landscapes: estuarine areas, mountain ranges, northern plains, intermediate valleys, and about 90 km of Pacific coastline. Elevations in the study area range from sea level to 1229 m at Cerro Ocotal, located within the Conchagua Volcano. The reference DEM for El Salvador is derived from airborne LiDAR (Light Detection and Ranging) data with a spatial resolution of 1 m and was provided by the Ministry of Environment and Natural Resources of El Salvador (MARN). The dataset is referenced to the NAD27 El Salvador datum, projected using a three-parameter Lambert Conformal Conic projection and based on the Clarke 1866 ellipsoid. The LiDAR data achieved a fundamental vertical accuracy of 0.178 m (95% confidence) in open terrain, while the final DEM reached a vertical accuracy of 0.202 m. A consolidated vertical accuracy of 0.182 m was obtained across all land cover types based on 480 ground control points, and horizontal accuracy was estimated at 1.118 m (95% confidence). Data acquisition was carried out using an ALTM 3100 airborne laser scanner (Teledyne Optech, Toronto, ON, Canada) with a ±20° scan angle, up to four returns per pulse, and a flight line overlap of approximately 60%, ensuring multiple-pass coverage and homogeneous point density (nominal pulse spacing of ~1.4 m).
Based on the LiDAR data, a normalized surface model (nDSM = DSM − DTM) was generated to stratify the surface into five height classes (
Figure 1d): bare ground (0–0.2 m), low vegetation (0.2–0.5 m), shrubs (0.5–2 m), trees (2–5 m), and tall trees (>5 m). The upper limit of the bare ground class was set using the Mean Vertical Error of the LiDAR, approximately 0.2 m, treating it in the nDSM as measurement noise. Buildings were not considered in the analysis, since, although the study area includes small urban clusters, the constructions are scattered and do not exceed two stories in height, making them largely irrelevant for land cover classification. Visually, the predominant vegetation in the area corresponds to heights between 2 and 5 m.
The chosen open access global models have been the focus of numerous previous studies, which facilitates contextualizing and comparing the results obtained in this work. Each model is based on different acquisition and processing techniques, directly influencing their accuracy and applicability in diverse topographic environments. The main characteristics are summarized below (
Table 1):
SRTM V3.0 is the highest-quality release of NASA’s Shuttle Radar Topography Mission DEM, providing a near-global digital surface model (DSM) at near-global 1 arc-s (~30 m) resolution. The dataset was collected in 2000 and has filled previous data gaps using sources such as ASTER GDEM2 and other auxiliary DEMs. Its vertical accuracy is approximately 9 m (90% confidence) [
39,
40,
41].
ASTER GDEM (Version 3)—Global DSM with a resolution of 1 arc-s (~30 m), based on stereoscopic imagery captured by the ASTER sensor aboard NASA’s Terra satellite. Released in 2019. The vertical/radial error may be as high as 20 m at the 95% confidence level [
42,
43].
ALOS World 3D (AW3D30)—Global DSM with 1 arc-s (~30 m) resolution, generated from optical stereoscopic images captured by the PRISM sensor onboard the ALOS satellite from 2006 to 2011. Released in 2014, it has a Root Mean Square Error (RMSE) of approximately 5 m [
44,
45].
Copernicus DEM (GLO-30)—Global DSM with a resolution of 1 arc-s (~30 m) with an average vertical accuracy of ~2 m, derived from radar data acquired during the TanDEM-X mission between 2010 and 2015. Released in 2019 [
46].
MERIT DEM—Three arc-s (~90 m) digital terrain model (DTM) published in 2017. It combines data from SRTM v3, AW3D, and other global models. Corrects systematic errors due to vegetation, buildings, and noise. Average vertical accuracy ~2 m [
47].
NASADEM Merged DEM Global 1 arc second V001—Global DSM at 1 arc-s (~30 m) resolution, released in 2020 by NASA’s LP DAAC. It improves the original SRTM data by integrating ICESat/GLAS lidar data, ASTER GDEM v2, ALOS PRISM AW3D30, and other sources to fill voids and reduce errors. NASADEM achieves an average vertical accuracy of approximately 2–3 m [
48].
FABDEM V1-2—One arc-s (~30 m) DTM based on Copernicus DSM, corrected using machine learning algorithms and LiDAR data from 12 countries. Improves accuracy in vegetated and urban areas [
49].
The evaluated DEMs use up to three different vertical reference systems. ASTER and SRTM, along with their derived models MERIT and NASADEM, are referenced to EGM96. ALOS AW3D30 uses orthometric heights relative to mean sea level, resulting in a model similar to a geoid equivalent to EGM96. Copernicus and its derived model FABDEM employ EGM2008. Since each datum has its own uncertainty, it was considered relevant to preserve these uncertainties in the calculation of the effective elevation differences, avoiding conversion processes that could introduce additional errors. Therefore, no datum transformations were applied. Accordingly, the native spatial resolution of each DEM was preserved throughout the workflow (1 m for LiDAR, 30 m for 1-arc-s DEMs, and 90 m for MERIT). Consequently, the elevation differences reported in this study reflect both acquisition-related errors, errors associated with spatial resolution, and discrepancies linked to the datum used, allowing the performance of each DEM to be evaluated under real use conditions.
To guarantee horizontal alignment with the LiDAR reference model, all DEMs were reprojected from WGS84 to NAD27 El Salvador using a three-parameter Lambert Conformal Conic projection. Nearest-neighbor resampling was applied during reprojection to preserve the original elevation values and the pixel connectivity characteristic of each DEM. A fixed grid origin coinciding with the LiDAR DEM was imposed to minimize subpixel misalignment between datasets. This reprojection, as well as the DEM–LiDAR comparison, was performed in QGIS Desktop software v3.34.5 (QGIS Development Team. Open Source Geospatial Foundation, Beaverton, OR, USA). Subsequent spatial and statistical analyses were conducted in Python v3.12.3 (Python Software Foundation, Wilmington, DE, USA) using the Rasterio (v1.4.3), NumPy (v1.26.4), and Matplotlib (v3.8.4) libraries.
Elevation comparisons were carried out using a raster-to-raster approach, taking the LiDAR cells as the reference. Each LiDAR pixel was paired with the elevation value of the corresponding overlapping DEM cell. This allows the error associated with the spatial resolution of each DEM to be incorporated into the calculation of effective differences.
2.2. Quantitative Evaluation of Global Error and Distribution Analysis
Beyond elevation, the analysis included slope, calculated as the magnitude of the elevation gradient, representing the maximum change in elevation between adjacent pixels.
Slope was calculated after reprojecting the DEMs to the same reference system as the LiDAR. Each derived model was compared using the same procedure applied for elevation, pairing each pixel of the 1 m LiDAR-derived model with the corresponding cell of the global derived model. Models were evaluated at their native resolution, preserving their original datum. Therefore, the effective differences in slope results from the combination of various sources of error: spatial resolution, acquisition characteristics of each DEM, and error propagation arising from the derivation of morphometric variables from elevation.
Comparing the variable derived from GDEMs with those obtained from the high-resolution LiDAR reference model allows for the identification of spatial patterns or inconsistencies and helps determine if systematic elevation errors propagate into the derived products. The LiDAR-based DEM is assumed to represent the most accurate available terrain surface and was used as the true value in elevation error assessments. The analysis was limited to pixels with valid values and elevations above 0 m in the reference LiDAR DEM, which was used to define the coastline. This approach excluded submerged areas and regions without data coverage. The decision is based on the variable capability of sensors to capture topographic information over water-covered surfaces, as well as on differences in sea level definitions across models, which result in coastline discrepancies and non-equivalent spatial extents.
Histograms were generated to analyze the distribution of elevation and slope values for each DEM and for the reference LiDAR model, always using relative frequencies to compensate for the lower pixel count of the MERIT DEM (90 m resolution). Global errors were quantified using pixel-by-pixel difference maps: DoD (Difference of DEMs) and DoDs (Difference of Slope Models). The following statistical metrics were derived from these maps to characterize the differences in elevation and slope.
Mean Error (ME)—ME represents the average of the direct differences between the evaluated DEM and the reference DEM. It allows the identification of systematic errors—such as consistent overestimations or underestimations—although it may conceal the presence of opposite-signed errors by canceling them out [
1,
6,
15,
16,
50].
where terms are defined as follows:
: Total number of observations;
Corresponding value from the LiDAR reference DEM;
Value of the evaluated DEM.
The error for each pixel was calculated as:
Mean Absolute Error (MAE)—MAE calculates the average of the absolute differences between the values of the evaluated DEM and the reference DEM. Unlike ME, it does not take the direction of the error into account, thus preventing the compensation between positive and negative errors [
4,
6,
51]
Root Mean Square Error (RMSE)—RMSE measures the average magnitude of the square deviations between observed and predicted values, giving greater weight to larger errors, which makes it sensitive to outliers. Furthermore, by squaring the differences, it also avoids the compensation between positive and negative errors [
1,
4,
6,
7,
15,
16,
50,
51].
2.3. Model Difference Assessment in Complex Coastal Areas and Proposal of the Spatial Error Robustness Index (SERI)
The analysis focuses on a 2 km strip inland from the coastline. Due to the wide geomorphological diversity of coastal environments and the variety of scientific and management objectives associated with coastal studies, there is no single or universally accepted definition for delineating the coastal zone. Therefore, the delimitation of this zone is usually adapted to the specific objectives of each study and the particular characteristics of the area analyzed. In this work, the coastal zone was defined as a 2 km inland buffer based on the purpose of evaluating DEM suitability for the identification and analysis of uplifted marine terraces within an active tectonic setting. This distance was intentionally selected to ensure that both the target landforms and their surrounding geomorphological context were adequately represented in the analysis.
Global statistical indicators can measure the overall error of the models, but they are insufficient to describe how errors are distributed spatially or to evaluate the geometric consistency between pixels. Moreover, traditional scatter plots become ineffective with large datasets due to overplotting and visual saturation, limiting their ability to reveal detailed error patterns. To address these limitations, DoD heatmaps were generated.
These plots group adjacent data points into hexagonal cells, encoding their density through a color scale, which effectively prevents visual saturation and overplotting. This approach provides a clear and detailed visualization of the distribution of differences between models across topographic variables. By plotting these values against corresponding reference values, it facilitates direct comparison along the gradient of each variable and enables the identification of spatial patterns or systematic biases that may be masked by global summary statistics or conventional distribution plots.
To quantify which models exhibit a better fit, the Spatial Error Robustness Index (SERI) was developed. This index enables a numerical interpretation of the distribution patterns observed in the heatmaps, facilitating the identification of the model that most closely aligns with the reference model used as a proxy for the topographic surface.
The SERI is inspired by the RSR (Root Mean Square Error–Standard Deviation Ratio), an index widely used to assess the performance of hydrological and environmental models [
52,
53,
54,
55,
56,
57], which measures the relative error of a model with respect to the natural variability of the observed data, with lower values indicating a better fit.
Unlike the original RSR, the calculation of SERI employs a modified RSR, in which the value range of the reference model is divided into equally sized intervals. For each interval, the RMSE and the standard deviation of the errors are calculated to derive a specific RMSE/SD error ratio, which is then weighted by the data density within that interval (i.e., the proportion of total data points it represents). These weighted ratios are subsequently summed to produce a global ratio that robustly characterizes the relationship between magnitude and natural variability of the errors across the elevation range, enabling the identification of whether errors are predominantly driven by systematic bias or random noise. Values tend to be equal to or greater than 1, with values near 1 indicating errors dominated by random noise without systematic bias, whereas values exceeding 1 indicate the presence of bias within that elevation range. Although values below 1 could theoretically occur in scenarios with high symmetrical dispersion and zero systematic bias, such cases are rare in practice.
However, the modified RSR alone may indicate the presence of bias but does not discriminate against the overall magnitude of the errors. Consequently, a model with large errors and high dispersion may yield a similar ratio to a model with small errors and low dispersion. For this reason, the SERI multiplies the weighted RMSE/SD values by two additional factors that penalize models with an RMSE above the mean of all evaluated models and/or with standard deviations above this mean. In this way, the SERI functions as a dimensionless, density-weighted index that simultaneously evaluates the magnitude and variability of errors, facilitating coherent comparisons among evaluated DEMs. Lower SERI values correspond to models that align more closely with the reference model, indicating a superior overall fit and higher geomorphological coherence. The calculation of SERI is given by the following equation:
: Root Mean Square Error for bin b.
: Standard deviation of errors for bin b.
: Weight for bin b, calculated as:
: Number of pixels inside bin b.
: Total number of pixels across all bins.
: Number of points in bin b.
: Mean of the errors for bin b.
Value of the model to be evaluated.
4. Discussion
4.1. Methodological Approaches
This study assessed the vertical accuracy of seven open access Digital Elevation Models (DEMs) against a high-resolution (1 m) LiDAR-based model. As noted by [
1,
2], the accuracy of DEMs is strongly influenced by both their spatial resolution and the specific conditions under which they were acquired. Therefore, no datum reprojections or spatial resampling were performed. This methodological decision aimed to preserve the original errors inherent to each DEM’s acquisition method and native resolution, ensuring that the differences were computed directly at the resolution of the reference model (1 m), without altering the spatial characteristics of the input DEMs.
The methodology employed is based on a combined approach that weights error metrics according to the density of observations, allowing simultaneous evaluation of both error magnitude and dispersion. The SERI values were consistent with the error patterns observed in the heatmaps for each model, indicating good performance of this index. Unlike traditional metrics, which are still widely used but do not provide information about the spatial distribution of errors, the SERI penalizes both extreme errors and variability, providing a more robust measure of geomorphological coherence between the evaluated model and the reference. Additionally, since it is constructed from common metrics (RMSE and standard deviation), it is easily interpretable, linear, and compatible with previous assessments.
4.2. Accuracy Evaluation by Variable
4.2.1. Elevation
All models exhibit a systematic overestimation, as evidenced by the positive mean error (ME) values across the entire study area and by the patterns observed in the coastal zone heatmaps. In particular, between 0 and 100 m of elevation, a clear positive bias can be observed, characterized by a higher density of positive errors, whereas higher elevation ranges exhibit a greater dispersion of errors. As reported in the literature [
15,
20], vegetation is one of the major sources of error in global DEMs, as most spaceborne optical and SAR sensors primarily interact with the vegetation canopy rather than the bare ground. This results in elevation surfaces that are closer to a DSM than to a true DTM and amplifies the positive bias in densely vegetated areas. This explains better ME and SERI obtained for models with systematic vegetation corrections, such as MERIT (ME = 1.62), and/or those fed with LiDAR data like FABDEM (ME = 2.62).
A clear correspondence exists between DEM performance and acquisition technique, vertical datum, spatial resolution, and post-processing strategy. Radar-derived DEMs generally outperform optical DEMs (ASTER and ALOS AW3D30) [
29], which are more prone to errors arising from vegetation, clouds, and sensor limitations [
7,
18,
20].
ASTER recorded the largest absolute error (106.79 m), attributed to residual artifacts despite applied corrections [
18], and shows the second-worst performance, as also reported by Bielski et al. [
4] and Guth et al. [
12]. ALOS exhibits better performance, but still shows limitations in densely vegetated areas, as optical sensors have lower penetration capability [
20,
29].
MERIT DEM, despite being designed to remove vegetation, buildings, and systematic noise from its parent datasets (ASTER and SRTM), shows the poorest performance for the entire area (RMSE = 9.94 m) and for the coastal area (SERI = 2.89). This indicates that its systematic correction does not guarantee improved accuracy when coarse spatial resolution (90 m) dominates the error structure, as also noted by Uuemaa et al. [
15].
NASADEM outperforms SRTM (the base model) in terms of elevation (RMSE = 7.19 m and 8.35 m for the entire area, and SERI = 1.01 and 1.41 for the coastal area, respectively). However, this improvement does not extend to slope, where SRTM performs similarly or even better. This suggests that reducing elevation errors through DEM editing, as implemented in NASADEM, does not automatically enhance the accuracy of derived terrain products such as slope [
12]. NASADEM also exceeds the Copernicus DEM (RMSE = 7.51 m for the entire area and SERI = 1.15 for the coastal area) in elevation accuracy, as the X-band radar used by Copernicus penetrates vegetation less effectively than C-band systems, leading to increased vertical errors in densely vegetated areas [
24,
58].
However, Copernicus, despite being a DSM without explicit vegetation removal, demonstrates robust and consistent performance in both elevation and slope [
18,
20]. FABDEM, derived from Copernicus, achieves the best elevation results for both the entire area (RMSE = 5.54 m; MAE = 3.74 m) and the coastal area (SERI = 0.81). This improvement over Copernicus is attributed to machine-learning-based corrections and the use of LiDAR data, which reduce biases associated with forests and buildings.
Several studies have agreed on selecting FABDEM as the most accurate Digital Elevation Model among other available DEMs. Gesch [
19] obtained the best statistical metrics for FABDEM in his study evaluating low-elevation coastal zones in the United States. Osama et al. [
59] selected it in their study using data from Haiti. Similarly, in non-coastal regions, FABDEM was also preferred by Huang & Yang [
60] in forested areas of North America, by Meadows et al. [
20] in a global study covering 65 flood-prone areas, and by Marsh et al. [
61] in a mountainous region of Canada. Finally, Bielski et al. [
4] conclude that Copernicus and FABDEM are the most accurate models in terms of elevation after being evaluated over 24 areas worldwide.
The superior performance of Copernicus and FABDEM is attributed to their radar-based acquisition and the use of the EGM2008 vertical datum, which integrates modern satellite missions along with denser terrestrial and marine gravity data. It provides a more accurate representation of regional and local gravitational variations than EGM96 [
62], with differences of up to 12 m reported in some regions [
63].
Nevertheless, despite the general consensus, several studies report important context-dependent exceptions. For example, Saberi et al. [
64] describe that MERIT DEM has outperformed FABDEM in Iran despite its coarser resolution. In extremely low-gradient areas, models such as SRTM and NASADEM sometimes exhibit slightly higher elevation accuracy than Copernicus [
12,
29]. Additionally, some studies indicate that the improvements introduced by FABDEM are not always statistically significant compared to Copernicus DEM [
59].
4.2.2. Slope
The slope analysis confirms that slope and elevation errors are not directly proportional, as they are distinct measurements. Slope, as the derivative of elevation, often exhibits amplified errors, particularly when comparing datasets with significantly different resolutions [
16,
18]. This is evident in coastal zones, where global DEMs tend to overestimate nearly flat terrain and very low slopes (<10°), but progressively underestimate slope as terrain steepness increases. The predominance of negative differences across most of the slope range is consistent with the negative ME values obtained for all models, indicating an overall tendency to underestimate slope throughout the study area. This reflects a systematic smoothing effect due to their coarser resolution compared to LiDAR.
Regarding model performance, Copernicus demonstrates the best slope accuracy in coastal zones (RMSE = 7.03°, MAE = 4.45°, SERI = 1.06), indicating good geomorphic connectivity despite not achieving the best elevation accuracy. FABDEM closely follows (RMSE = 7.15°, MAE = 4.48°, SERI = 1.09). In contrast, ALOS shows the best RMSE and MAE for slope across the entire study area but the worst SERI in coastal zones (1.51). This aligns with previous findings [
12,
15,
17,
20] which indicate that while ALOS demonstrates superior performance in steep terrain, its accuracy may be comparatively reduced in low-relief environments such as coastal zones.
These results highlight the need to assess DEMs within specific geomorphological contexts rather than depending exclusively on global statistics, as model accuracy varies with terrain and scale. Performance in elevation does not necessarily predict performance in derived parameters like slope.
This pattern has been observed in other studies. Osama et al. [
59] found that while FABDEM and Copernicus achieved high elevation accuracy, ALOS outperformed them in slope estimation. Marsh et al. [
61] reported slight slope errors in FABDEM on steep slopes, and Santillan [
65] observed reduced accuracy in areas with slopes > 2° or elevations > 100 m. Guth et al. [
12] emphasized that although AI-enhanced DEMs like FABDEM improve elevation accuracy, this does not necessarily translate into improved slope precision, a pattern confirmed in our study where Copernicus slightly outperforms FABDEM for slope estimation in coastal zones.
4.2.3. Best Model Selected
Among the evaluated DEMs, NASADEM, FABDEM, and Copernicus stand out for their overall superior performance in both elevation and slope (
Figure 6). All three were acquired using radar SAR sensors. Copernicus and FABDEM use the EGM2008 vertical datum, while NASADEM and FABDEM incorporate LiDAR data during post-processing. To visually assess which model outperforms the others,
Figure 7 displays the per-pixel error maps in the coastal zone, using elevation as a reference and highlighting areas of higher geomorphic relevance.
While the overall error distributions among these models are similar, FABDEM improves precision in estuarine areas (zones a and c in
Figure 7) and reduces noise in cliffs and high-relief regions (zone b in
Figure 7) compared to Copernicus and NASADEM. FABDEM stands out as the most accurate option among available global DEMs in the coastal zones. Several recent studies [
4,
20,
59,
61,
66] have confirmed the improved accuracy of FABDEM over Copernicus in vegetated and inundated terrains. Additionally, Guth et al. [
12] noted that while FABDEM improves over Copernicus in flat coastal areas, it does not necessarily outperform it in steep, rugged terrain.
Notably, contrary to expectations, zone b (steep cliffs) exhibits lower errors than estuarine zones, where errors predominantly exceed 30 m. Consistent with these observations, Li et al. [
29] found that steep, vegetation-free areas in China had lower errors than tall, densely vegetated regions, while Meadows et al. [
20] emphasized that land cover is the primary factor influencing vertical DEM errors, particularly in vegetated environments.
Although there is no specific information on FABDEM’s effectiveness in removing vegetation artifacts in our study area, comparison with Copernicus, the model on which it is based, suggests that these corrections reduce vegetation-induced elevation errors from over 30 m to (−5)–5 m (
Figure 7). This indicates that, even without site-specific validation, FABDEM significantly improves vertical accuracy in complex, densely vegetated terrain, contributing to a more reliable representation of coastal and volcanic topography.
4.3. Multi-Scale Sensitivity of DEM Performance and SERI Robustness
To verify the stability of the results obtained within the 2 km coastal zone and evaluate their transferability to more complex geomorphological settings, RMSE values, SERI indices, and error distributions were recalculated using 4 km and 10 km inland buffers measured from the coastline. This multiscale comparison allowed the evaluation of whether the relative performance of the DEMs remained consistent as the analysis domain expanded, while also examining the sensitivity of elevation and its derived topographic variables to increasingly heterogeneous geomorphological conditions.
4.3.1. Multi-Scale Elevation
The results indicate that the overall ranking of the DEMs in terms of elevation accuracy remains remarkably stable as the spatial extent of the analysis increases. Although the 4 km and 10 km buffers incorporate greater geomorphological complexity, including higher elevations, volcanic edifices, steeper terrain gradients, and more deeply incised drainage networks, the highest- and lowest-performing DEMs remain unchanged across all spatial scales considered.
FABDEM consistently yielded the lowest SERI and RMSE values across all buffer extents, demonstrating superior geomorphological consistency and vertical accuracy, while MERIT DEM exhibited the highest SERI and RMSE values (
Table 4). DEMs exhibiting intermediate performance experienced only minor shifts in their relative rankings between buffer extents. These variations are more strongly related to changes in the spatial organization and distribution of errors than to substantial differences in their overall magnitude.
For instance, within the 4 km buffer, ASTER outperformed ALOS despite displaying a broader but largely symmetric error distribution, whereas ALOS exhibited lower dispersion but maintained a persistent positive bias across a substantial portion of the elevation range. However, when the analysis was extended to 10 km, ASTER showed a more pronounced positive bias, while ALOS improved its relative position due to reduced error dispersion, even surpassing Copernicus DEM.
In the 10 km buffer analysis, NASADEM and ASTER exhibited positive anomalies (overestimation) concentrated approximately between 250 and 400 m elevation (
Figure S2 in the Supplementary Material). The spatial distribution of these errors coincides with sectors associated with incised ravines developed on volcanic slopes, suggesting limitations in the representation of narrow linear landforms and abrupt topographic discontinuities. Conversely, NASADEM and SRTM presented negative anomalies (underestimation) around 250 m elevation, mainly associated with ridgelines and drainage divides characterized by pronounced curvature gradients. The spatial correspondence between these residual patterns and areas of high geomorphological variability indicates that NASADEM, SRTM, and ASTER are more prone to topographic smoothing, attenuating sharp terrain transitions. Although these limitations are inherent to all global DEMs due to their spatial resolution, they are particularly evident in these three datasets.
The overall increase in RMSE values with increasing distance from the coastline (4 km and 10 km) suggests that altimetric discrepancies are amplified as more complex landforms are incorporated. Nevertheless, the overall ranking of the DEMs remains remarkably stable, highlighting that their relative altimetric performance is largely insensitive to the spatial extent of the study area.
4.3.2. Multi-Scale Slope
Unlike elevation, the results obtained for slope show greater sensitivity to the extent of the analyzed area. The relative ranking of the models varies more substantially, indicating that the representation of terrain derivatives depends more strongly on the inter-pixel relationships within each DEM (i.e., how error propagates through neighborhood-based analyses used to calculate derivatives).
A particularly significant observation is the modification of the error distribution structure between the 2 km buffer and the larger 4 km and 10 km buffers. In the latter cases, the highest density of observations is concentrated within an approximately horizontal band ranging from +10° to −20° across most of the slope domain (
Figures S3 and S4 in the Supplementary Material). However, the distribution is distinctly asymmetric, extending from approximately −70° to +30°, evidencing a predominance of negative residuals and therefore a systematic tendency for global DEMs to underestimate slope gradients throughout the analyzed range.
Copernicus DEM and FABDEM show the best results within the 2 km buffer, while Copernicus performs best at 4 km (
Table 5). In contrast, MERIT DEM consistently records the highest SERI and RMSE values in the 4 km and 10 km buffers, confirming that the generalized terrain smoothing inherent to this dataset adversely affects the representation of terrain derivatives. ALOS exhibited the most notable change in performance from the lowest-ranked model at 2 km to the second-best performer at 4 km and the highest-ranked DEM at 10 km, progressively improving as the analysis distance increases. This result is consistent with the overall RMSE results. The observed pattern shows that ALOS represents the topographic gradients present in steep, high-relief terrain more effectively than the other DEMs evaluated, whereas its performance is less favorable in low-slope coastal environments.
The reduction in RMSE values observed between the 2 km buffer and the broader inland analyses indicates that the largest slope errors are concentrated in low-gradient coastal terrain. This reflects the intrinsic sensitivity of slope calculations to small vertical discrepancies in relatively flat terrain, where minor elevation errors can generate disproportionately large relative errors in derived slope estimates. As the proportion of moderate-to-high-slope terrain increases within the analysis area, the influence of these local errors decreases and global RMSE values tend to stabilize.
The results obtained using SERI and the spatial distribution analysis of errors for the different analysis buffers confirm that the altimetric assessment is highly robust to changes in coastal zone delineation. In contrast, slope assessment shows a greater dependence on both the geomorphological context and the intrinsic characteristics of each DEM. Therefore, the stability observed for elevation accuracy cannot be directly extrapolated to terrain derivatives such as slope, whose behavior is conditioned by error propagation mechanisms and by the spatial distribution of landforms. These findings emphasize that global error metrics alone may not adequately characterize DEM performance in specific applications and highlight the importance of incorporating analyses of error distribution and spatial error structure into DEM validation frameworks.
To further evaluate the consistency of the proposed SERI with conventional accuracy metrics, correlation analyses were performed between SERI, RMSE, and the standard deviation of errors (STD) across all DEMs and buffer extents considered in this section. Because SERI is derived from both error magnitude and error variability, relationships with RMSE and STD were evaluated independently (
Tables S1–S6, Supplementary Data).
Pearson correlation analyses (
Table 6) revealed strong positive relationships between SERI and RMSE, with coefficients consistently above 0.97 across buffer scales in the elevation-based analysis, indicating that the proposed index is highly sensitive to variations in overall error magnitude. Similarly, strong correlations were observed between SERI and STD, suggesting that the index also captures information related to error dispersion and spatial variability.
A distinct behavior was observed in slope-dependent analyses at the 2 km scale. In this case, the relationship between RMSE and STD became strongly negative, and the correlations between SERI and RMSE, as well as between SERI and STD, were notably reduced. This behavior is likely related to an inversion in the relationship between error magnitude and dispersion, where higher slope values are associated with more systematic (less variable) errors. In this context, SERI appears to penalize structured error patterns differently, reflecting its sensitivity to the spatial organization of errors rather than solely to their magnitude.
These results suggest that, under certain topographic conditions, a high level of consistency between SERI and traditional metrics may emerge when both describe a coherent relationship between error magnitude and dispersion. However, in other cases, the reduction in these correlations indicates that each metric may capture different aspects of error behavior. Consequently, the exclusive use of traditional metrics may not fully represent the complexity of error structure in heterogeneous topographic settings.
Overall, these findings highlight the value of SERI as an integrative metric, capable of simultaneously incorporating information on both error magnitude and spatial structure.
4.4. External Validation of FABDEM and Robustness of the Methodological Approach
To support the selection of FABDEM, a literature review was conducted where FABDEM reported the best performance according to traditional error metrics. Additionally, to strengthen the choice of this model and validate the proposed methodology for coastal areas, an external verification was carried out in geodynamically comparable areas to the target of future analysis—coastal margins in subduction zone contexts aimed at identifying uplifted marine terraces. A summary table (
Table 7) is presented below, showing the RMSE, MAE, and bias values reported for FABDEM in various global studies where it has been identified as the most accurate elevation model.
The results reported in other studies confirm that FABDEM consistently performs well in coastal and vegetated environments—such as the one examined in this study—typically achieving RMSE values below 6 m. While its accuracy may decrease slightly in mountainous areas or regions with dense vegetation, it still outperforms other freely available global DEMs.
As previously discussed, standard error metrics (e.g., RMSE, MAE, bias) provide a broad measure of vertical accuracy but do not necessarily capture the preservation of subtle geomorphological features, such as marine terraces.
The identification of terraces depends on the preservation of local slope changes, flat surfaces, and continuity of curvature. Satellite optical and radar sensors frequently capture elevations closer to the canopy top or urban structures rather than the bare ground, resulting in a systematic positive bias, as evidenced in
Figure 1d and in the LiDAR–GDEM comparison results shown in
Figure 4 and
Figure 5. This leads to smoothed or displaced terrace boundaries [
6,
20,
66], an effect that is exacerbated in low-slope or coastal areas, where elevation differences are typically more subtle.
The inter-pixel inconsistency introduced by vegetation or buildings propagates to DEM derivatives, producing unstable and noisy values. This complicates the automatic detection of knickpoints and geomorphological nodes to delineate terrace edges [
12,
18]. Similarly, slope calculations are also affected by artificial roughness, leading to an overestimation of steep slopes and an underrepresentation of flat or gently inclined terrace surfaces, thereby distorting terrace morphology [
18,
59].
To assess the transferability of the methodological approach and the robust performance of FABDEM in future studies applied to morphotectonics, the analysis was applied to three tectonically active coastal zones: the Bōsō Peninsula (Japan), Newport (Oregon, USA), and Turakirae Head (New Zealand). These regions were selected due to the availability of high-resolution LiDAR data and well-documented landforms, such as marine terraces uplifted by seismic activity [
67,
68,
69].
In all three sites, the reprojected FABDEM model was compared against local high-resolution datasets. The modified RSR values (RMSE/SD ratio) obtained were consistent with those from the coastal area studied in El Salvador, even showing better results, ranging from 1.01 to 1.07. These values, being so close to 1, indicate that the FABDEM model in these three areas does not exhibit systematic errors, and any observed dispersion can be attributed to random variation. This conclusion is further supported by the lack of notable systematic biases in the heatmaps (
Figure 8j–l), where errors in all three cases are centered around y = 0 and rarely exceed ±30 m. However, at the Newport site, the FABDEM heatmap (
Figure 8l) shows greater dispersion and a mild overestimation of elevation, which could be attributed to the fact that this site covers a wider range of elevation compared to the coastal strips in El Salvador, Boso, and Turakirae.
Moreover, the slopes derived from FABDEM (
Figure 8d–f) allowed for the identification of the lateral continuity of previously mapped marine terraces (
Figure 8a–c), reinforcing their usefulness for morphotectonic analysis. The longitudinal profiles (
Figure 8g–i), which compare the same transect line between the reference model and the evaluated models, provide a smoothed visualization of these terraces, further confirming the morphological features observed in the slope maps. This supports the conclusion that, although FABDEM was not the top-performing model for slope estimation, it is adequate for detecting morphological changes at the required scale. These findings validate the robustness of the workflow—combining heatmaps and the SERI—and confirm its reliability for application in other active coastal settings.
4.5. Limitations, Implications, and Recommendations
Although the results obtained in this study are robust, they should be interpreted considering certain inherent limitations related to the study design and the data used.
First, as with any Digital Elevation Model evaluation, accuracy was assessed against a reference model assumed to be ideal (LiDAR). However, LiDAR data inherently carry their own unquantified uncertainties, which were not accounted for in this study and may introduce systematic or random biases that remain uncorrected. While these biases are generally acceptable in comparative studies, they should be acknowledged. We did not calculate the precision of our reference data, as we did not have independent ground control points available for validation. As in any comparative evaluation, accuracy is analyzed relative to the reference model, not necessarily representing the absolute truth of the terrain.
Second, there is no standardized methodology for accuracy assessment in DEMs that accounts for the spatial distribution of error. Although most studies use global metrics such as RMSE, MAE, or ME, which facilitate comparison between studies, these do not consider the spatial distribution of errors nor their propagation into derivatives such as slope or curvature. In this regard, the proposed SERI combined with the heatmaps of error distribution represents a significant methodological contribution, while maintaining the interpretability of traditional metrics. Nevertheless, its implementation requires high-accuracy reference data, which may limit its applicability in resource-limited environments or where LiDAR coverage is lacking.
Third, the study was conducted in a coastal area of El Salvador, so the results should be interpreted as representative of its specific geomorphological, climatic, and land-cover conditions. The complementary evaluation of the best-performing DEM (FABDEM) in three additional coastal settings at geographically distinct locations supports the reproducibility of the methodology and confirms the suitability of FABDEM as a precise model for identifying subtle morphological changes.
This delimitation restricts the extrapolation of results to inland or mountainous areas, where the performance of models—even FABDEM—could be affected [
65,
66]. In this context, some studies [
15,
17] agree that DEM accuracy varies depending on local geographic characteristics (slope, vegetation, relief), so local validation should always be highly recommended before applying any global DEM to tectonic, environmental, or risk assessments. Rather than providing a universal selection rule, the proposed framework is useful for comparing DEM performance within heterogeneous landscapes and for identifying which model best fits the specific characteristics of the study area. Nevertheless, its implementation depends on the availability of high-precision reference data.
Fourth, the possible influence of terrain changes between the acquisition dates of the different evaluated DEMs was not specifically analyzed in this study. Since the main objective of this work is to assess the overall performance of each DEM under real application conditions, these possible variations are considered part of the total observed error rather than something to be analyzed in isolation.
In this sense, a more recent DEM will not necessarily provide a more accurate representation of the terrain if its acquisition or processing introduces other sources of error, such as insufficient vegetation removal or systematic biases associated with model generation. Therefore, even if some models do not capture certain topographic changes that occurred between different acquisition dates, their ability to coherently represent terrain morphology remains a valid criterion for evaluating practical performance. Likewise, when such changes are spatially localized, their potential influence can be identified through spatial error analysis and the SERI, allowing their relative relevance to be assessed within the study area as a whole.
Finally, although the proposed approach—combining heatmaps and SERI—has proven useful for revealing error patterns not captured by global metrics, SERI should be interpreted as a relative metric dependent on the analyzed dataset, rather than as an absolute performance indicator. In this context, other indices and metrics could also be used to assess the spatial coherence. These include measures of horizontal connectivity or morphometric indices that evaluate spatial coherence beyond simple dispersion. Additionally, higher-resolution DEMs (such as TanDEM-X) were not included because they are neither free nor globally available.
Effect of Coordinate System Selection
The comparison between results obtained in the projected metric system and those derived in WGS84 shows that the choice of coordinate reference system has only a minor effect on the overall evaluation of DEM accuracy in the study area. For elevation differences, all models exhibit very similar error magnitudes in both approaches, with variations generally limited to a few decimeters and without altering the main ranking of performance. FABDEM remains the best-performing model in both systems, followed by NASADEM and Copernicus DSM (
Table 2 in the main manuscript;
Table S7 in the Supplementary Material). The only noticeable change is observed for MERIT DEM, which improves more substantially in WGS84 and slightly moves ahead of ASTER GDEM V3 and ALOS-AW3D30 in the ranking based on elevation error metrics.
In the coastal subset, the analysis based on the RMSE/STD ratio and the SERI shows the same general pattern (
Figure 5 in the main manuscript;
Figure S3 in the Supplementary Material). Although the numerical values vary slightly between the projected and geographic systems, the relative performance of the DEMs remains essentially unchanged, confirming that the selection of the most suitable DEM for this region is robust to the coordinate system used for error calculation. Again, MERIT DEM stands out as the model most affected by the change in reference system, but its overall position remains poor because of its high error dispersion and low geomorphic coherence. This behavior may be related to its coarser native resolution of 3 arc-s, which makes it more sensitive to reprojection effects and, consequently, to subsequent processing steps after transformation, such as derived products like slope [
15].
These results suggest that, for this low-latitude study area, reprojection-related distortions are small enough not to modify the main conclusions regarding DEM suitability.
From a methodological point of view, these findings indicate that the system used to compute elevation differences does not materially affect the choice of the best DEM for this specific region, likely because the study area is close to the equator and the distortions introduced by reprojection remain limited. However, this should not be interpreted as evidence that the coordinate system is irrelevant. On the contrary, when the objective is to assess the intrinsic accuracy of a given DEM, even moderate reprojection effects may alter the error structure and, in some cases, reduce the comparability of the results. For this reason, the present analysis is useful not only to verify the stability of the DEM ranking but also to highlight that reprojection can introduce additional uncertainty that may be non-negligible depending on the model’s native resolution and geometric characteristics.
5. Conclusions
This research provides a comprehensive evaluation of seven freely available Digital Elevation Models (DEMs) using a 1 m LiDAR dataset as a reference. The analysis preserves the native characteristics of each DEM and combines conventional error metrics with spatial analyses and the newly proposed SERI. Validation was conducted in tectonically active coastal regions. The main conclusions are as follows:
FABDEM emerged as the most accurate DEM for elevation in both coastal and broader regional contexts (RMSE = 5.54 m; MAE = 3.74 m; SERI = 0.8). MERIT showed the lowest bias but performed poorly due to lower resolution, while ASTER and ALOS displayed higher errors linked to stereoscopic image limitations in coastal areas.
Slope errors were systematic across all models, amplified by scale differences. DEMs tended to overestimate gentle slopes (<15°) and underestimate moderate to steep slopes (>15°). ALOS achieved the lowest RMSE for slope at the regional scale, while FABDEM and Copernicus performed best in coastal zones.
The proposed SERI effectively complements traditional metrics by penalizing both error magnitude and spatial variability. Combined with heatmaps, SERI enabled the differentiation of systematic and random error patterns, improving geomorphic coherence assessments across DEMs.
Additional analyses performed in the original WGS84 coordinate system showed that reprojection into a projected metric system introduces only minor variations in elevation accuracy metrics. For this low-latitude study area, these variations do not substantially affect the overall DEM performance ranking. However, some DEMs, particularly MERIT DEM, exhibited greater sensitivity to reprojection effects. These findings highlight the importance of considering coordinate reference systems and reprojection-related uncertainty in DEM accuracy assessments.
Findings confirm that DEM accuracy depends on both the model and the geomorphic variable assessed. Careful DEM selection based on study objectives and terrain characteristics is essential. FABDEM, validated across diverse coastal and tectonic settings, offers a reliable basis for geomorphological analyses.