1. Introduction
The Brazilian Cerrado covers approximately 2 million km
2, constituting South America’s second-largest biome and a global biodiversity hotspot with more than 12,000 cataloged plant species and high endemism rates, including approximately 4800 species occurring exclusively in this region [
1,
2,
3,
4]. This tropical savanna exhibits a complex mosaic of phytophysiognomies, ranging from open grasslands to dense forests, creating complex ecological gradients. In addition to its ecological value, the Cerrado contains the headwaters of eight major South American hydrographic basins, with deep-rooted vegetation that maintains critical hydrological functions [
5]. The biome also comprises extensive tracts of secondary vegetation. Secondary vegetation can be understood as the natural regrowth of vegetation after an anthropogenic disturbance [
6]. It is increasingly recognized for its ecological importance in carbon sequestration, biodiversity conservation, and landscape connectivity restoration [
7]. Despite its ecological importance, in Cerrado, protected areas represent only 8.3% of the biome, and this percentage drops to 6.5% when considering only the fraction still covered by native vegetation [
8,
9].
The Cerrado is also a major agricultural expansion frontier in Brazil. Nearly half of the original vegetation has been converted mainly to croplands and pastureland in the last 50 years [
2]. In fact, deforestation in the Cerrado has been similar to or greater than that occurring in the Amazon in recent years [
10]. Furthermore, the Cerrado is vulnerable to deforestation and is currently excluded from major agribusiness sustainability efforts, receiving less protection from national and international regulations [
11]. This scenario requires effective mapping and monitoring of land use and land cover (LULC) to support Brazil’s environmental commitments, including Nationally Determined Contributions under the Paris Agreement and the restoration targets established by the National Plan for Native Vegetation Recovery (Planaveg) [
3,
12].
TerraClass is a project carried out by the National Institute for Space Research (INPE) and the Brazilian Agricultural Research Corporation (EMBRAPA) that produces LULC maps for Cerrado and Amazon. It performs thematic mapping of previously deforested areas [
6,
13]. In the specific case of secondary vegetation in the Cerrado biome, mapping is conducted manually through visual interpretation of spectral and spatial patterns observed in historical satellite image series. Although this method provides high thematic quality, it is time-consuming and requires intensive manual effort, which directly affects operational efficiency.
Recent advances in Earth observation satellites have improved LULC monitoring by providing multispectral imagery at finer spatial and temporal resolutions [
14]. The large datasets of images produced by these satellites allow the use of satellite image time series (SITS) for continuous land monitoring, capturing phenological cycles, gradual vegetation transitions, and LULC changes often missed by snapshot-based approaches [
15]. SITS combined with machine learning methods have shown great potential for large-scale LULC mapping [
16,
17,
18,
19]. Recent advances in this field have been driven primarily by supervised approaches that require a large number of high-quality training samples. For example, Simoes et al. [
19] used 48,850 training samples to map the Cerrado biome using SITS, while Alencar et al. [
20] mapped three decades of native vegetation change, based on annual Landsat composites, using up to 50,000 samples per tile filtered by temporal consistency criteria.
Although effective, supervised approaches present a significant challenge, as the size, quality, and representativeness of the training sample directly affect model performance and mapping accuracy [
21]. These challenges are amplified in regions with high temporal variability and heterogeneous landscapes, where obtaining training samples that adequately represent the diverse environmental conditions is particularly time-consuming. Therefore, assembling training datasets that adequately capture this diversity remains one of the main limitations of large-scale LULC mapping using supervised classification of remote sensing images [
22]. As an alternative to supervised methods, clustering-based unsupervised approaches are promising options, as they exploit intrinsic SITS patterns without requiring labeled training data. As demonstrated by Shahi et al. [
23], temporal clustering of Sentinel-2 SITS can effectively distinguish different vegetation types, reducing labeling costs and dependence on reference data.
Self-organizing maps (SOM) [
24] combined with hierarchical clustering analysis (HCA) [
25] have been used in remote sensing to group spectral variability in an unsupervised manner. Gonçalves et al. [
26,
27] applied this SOM-HCA framework to single-date LULC classification. SOM-based clustering has been used in environmental studies to group landscape units and relate them to LULC patterns [
28,
29]. In the context of SITS, Santos et al. [
30] combined SOM and HCA to identify spatiotemporal patterns and characterize intra-class variability in LULC samples based on spectral and phenological behavior. Dynamic Time Warping (DTW) [
31] has been established as a suitable distance metric for comparing temporal trajectories with differences in timing, phase, or phenological alignment [
17,
32].
In alignment with this methodological direction, we propose a clustering-based methodology for large-scale LULC mapping that integrates SITS with unsupervised learning and expert knowledge. Our approach uses SOM followed by HCA, with DTW as the distance metric, to map secondary vegetation across the Cerrado biome. This design reduces the need for extensive training datasets by shifting expert interpretation from pixels to clusters, lowering manual effort while preserving thematic quality. The proposed methodology was evaluated for secondary vegetation mapping across the Cerrado biome within the TerraClass project.
3. Results
Due to the TerraClass methodology (
Section 2.2), which defines deforestation only within areas mapped by PRODES up to the reference year, the study area, after masking out-of-scope areas (
Section 2.3.4), comprises 691,958.83 km
2 of deforested land in the Brazilian Cerrado. Within this domain, the semi-automated workflow mapped 81,209 km
2 as secondary vegetation, corresponding to 11.74% of the study area (
Figure 5) (Area calculated from the final product after reprojection to SIRGAS 2000 (EPSG:4674), the coordinate reference system of the official TerraClass Cerrado 2024 release; it differs by about 11 km
2 (∼0.014%) from the value computed in the BDC Albers Equal-Area Conic projection used throughout processing, a reprojection artifact at polygon boundaries that does not reflect a calculation error). The ecoregions with the largest areas of secondary vegetation were the Floresta de Cocais and Paraná Guimarães, with 13,380 km
2 and 12,730 km
2, respectively. On the other hand, the ecoregions with the smallest areas of secondary vegetation were the Depressão Cárstica do São Francisco (374 km
2) and Costeiro (589 km
2) (
Supplementary Materials, Table S2).
3.1. Map Refinement and Composition
The seven-component framework results are presented from two complementary perspectives.
Figure 6 shows the processing perspective, in which all seven components are expressed as a proportion of the total secondary vegetation processed (SV processed) within each ecoregion (
Supplementary Materials, Table S1). SV processed encompasses every area that entered the pipeline as secondary vegetation at any stage, including areas from the initial cluster labeling and areas subsequently added through post-processing or manual editing. Because the seven components partition this total, they sum to 100% for each ecoregion. The concordance share directly measures the fraction of SV processed that required no modification, providing a per-ecoregion indicator of editing intensity: the higher the concordance, the less intervention during refinement was needed.
In this perspective, concordance rates ranged from 77.2% in Floresta de Cocais to 33.9% in Planalto Central (
Supplementary Materials, Table S1). Planalto Central also had the highest post-processing removal rate (33.8%). The highest manual addition shares relative to SV processed were observed in Chapadão do São Francisco (31.3%), followed by Complexo Bodoquena (19.1%). Basaltos do Paraná and Complexo Bodoquena had the highest recovery rates, 6.4% and 6.9%, respectively.
The composition perspective (
Supplementary Materials, Table S2) shows where the final secondary vegetation map originated from. From this perspective, expressed as a share of the final SV area, concordance accounts for 95.0% of the final map in Alto Parnaíba but only 48.6% in Chapadão do São Francisco. Ecoregion-level redundancy details are provided in
Supplementary Materials, Table S3.
Cluster labeling initially identified 120,398 km
2 as secondary vegetation (
Table 4). Considering only pixels that originated in cluster labeling, automated post-processing removed 29,801 km
2 (24.8%) via MMU spatial filtering (<2 ha), and manual review excluded an additional 20,519 km
2 (17.0%), totaling 50,320 km
2 (41.8%). The remaining 70,078 km
2, referred to as
concordance, corresponds to the secondary vegetation retained after refinement and represents 58.2% of the initially labeled secondary vegetation. In the final composition (
Table 5), concordance is reported as 70,076 km
2, slightly lower than the 70,078 km
2 reported above, reflecting the application of the final agriculture mask after refinement.
The final product also includes areas added during post-processing and manual editing.
Table 5 decomposes the final secondary vegetation area by source. Concordant areas account for 86.3% of the final secondary vegetation. Automated gap filling contributed 4111 km
2 (5.1%), manual editing added 5377 km
2 (6.6%), and 1656 km
2 (2.0%) represents areas removed by automated filtering but later reinstated by interpreters.
Table 5 was calculated in the BDC Albers Equal-Area Conic projection, whereas the final product was reported in SIRGAS 2000 (EPSG:4674); the 11 km
2 (∼0.014%) difference is a boundary reprojection artifact.
As described in
Section 2.2, the agriculture mask from the current TerraClass cycle is applied after the secondary vegetation refinement stage.
Supplementary Table (Supplementary Materials, Table S3) summarizes all final removals during refinement, including removals from both cluster-origin and gap-fill-origin pixels. Of this total 52,209 km
2 removed during refinement, 4527 km
2 (8.7%) overlapped with the agriculture mask (
Supplementary Materials, Table S3). These removals are termed
redundant, as the corresponding areas would have been masked regardless of editorial action. The remaining 47,683 km
2 (91.3%) constitute effective removals. Redundancy rates varied across ecoregions, from 0.2% in Depressão Cárstica do São Francisco to 26.6% in Basaltos do Paraná.
The ecoregion-level composition rates (
Supplementary Materials, Table S2) show that concordance in the final map ranged from 95.0% in Alto Parnaíba to 48.6% in Chapadão do São Francisco, with recovery peaking in Basaltos do Paraná and manual addition in Chapadão do São Francisco.
3.2. Accuracy Assessment
Accuracy was assessed through stratified random sampling following Olofsson et al. [
59], with 695 validation points (241 secondary vegetation, 454 pasture/background by map stratum).
Table 6 presents the area-adjusted accuracy metrics for both classes and the overall accuracy. For secondary vegetation, the final product achieved a user’s accuracy of 96.27% and a producer’s accuracy of 79.22%. The corresponding F1-score for secondary vegetation is 86.90%. The estimated area for secondary vegetation was 98,683 km
2 ± 10,071 km
2 (95% confidence interval (CI): 88,612–108,754 km
2). The mapped area (81,209 km
2) falls below the lower bound of this interval, indicating systematic underestimation relative to the reference data. The overall accuracy was 96.45%. At the point level, validation points yielded 16 omission errors and 9 commission errors.
4. Discussion
Mapping secondary vegetation is challenging due to its spatial characteristics (fragmentation) and age differences [
6,
20]. The concordance and composition results (
Table 4 and
Table 5) indicate that SOM-HCA captures the main secondary vegetation spatial signal while tolerating overestimation at the clustering stage. Most of the final product originates from clusters initially labeled as secondary vegetation (86.3%), and refinement is therefore predominantly subtractive. This behavior is consistent with a labeling strategy that prioritizes coverage and delegates class separation to the expert stage.
4.1. Secondary Vegetation Refinement
A large share of the initial overestimation is structural and reflects the 2-ha MMU constraint. The spatial filter removes all fragments smaller than 2 ha (24.8%;
Table 4). This effect is not uniform across the Cerrado. It depends on landscape structure, because transition zones often split secondary vegetation into small fragments that fall below 2 ha. The different values across ecoregions show this heterogeneity through differences in post-processing removal and recovery (
Supplementary Materials, Table S1;
Figure 6). The recovered component further indicates that some MMU-driven removals correspond to true secondary vegetation fragments that are later recovered by interpreters (
Table 5).
Beyond MMU effects, residual overestimation reflects spectro-temporal mixing within clusters. Cluster homogeneity varies with landscape complexity and the sharpness of vegetation transitions. Under gradual Cerrado transitions, secondary vegetation-labeled clusters may include nearby non-secondary vegetation pixels, and imposing a stricter decision boundary would also remove true secondary vegetation along margins and within heterogeneous mosaics. This component is addressed during expert screening and is consistent with the share removed manually (17.0%;
Table 4).
Part of the refinement effort also targeted areas later masked as agriculture, because both products were mapped concurrently (
Supplementary Materials, Table S3). This overlap and its implications are discussed in
Section 4.7. The subtractive nature of this refinement strategy has direct consequences for the accuracy of the final product.
4.2. Ecoregion Variability
The validation design was stratified at biome scale, so the ecoregion-level values reported here should be interpreted as processing indicators rather than accuracy estimates with formal confidence intervals.
Ecoregion-level contrasts indicate that the same pipeline interacts differently with distinct landscape structures, so performance is better interpreted at the ecoregion level than through a single biome-wide metric. The processing perspective (
Supplementary Materials, Table S1;
Figure 6) reflects editing burden, whereas the composition perspective (
Supplementary Materials, Table S2) reflects how much of the final map still derives from the initial clusters.
This distinction helps explain why different ecoregions stand out depending on the metric considered. Floresta de Cocais, for example, showed the highest concordance in the processing perspective (77.2%), indicating that the initial clustering already matched the final product closely. Alto Parnaíba, in turn, showed the highest dependence of the final map on the initial clusters (95.0%), with little need for manual addition. By contrast, Planalto Central and Basaltos do Paraná were more strongly affected by post-processing removal and recovery, consistent with fragmented landscapes in which the 2-ha MMU has greater influence. Chapadão do São Francisco represents a different limitation, where underdetection in the initial clustering was later compensated by manual addition. Overall, these patterns align with the landscape heterogeneity and transition complexity previously reported for the Cerrado [
36].
4.3. Accuracy Analysis
For secondary vegetation, the final product achieved a user’s accuracy of 96.27% and a producer’s accuracy of 79.22% (
Table 6). This asymmetry is a direct consequence of the subtractive refinement strategy: the high user’s accuracy indicates that mapped secondary vegetation areas are reliable, while the lower producer’s accuracy reflects that a substantial portion of the reference extent was not captured. This is corroborated by the area estimates: the mapped area (81,209 km
2) falls below the lower bound of the 95% confidence interval for the estimated area (88,612–108,754 km
2;
Table 6). The primary source of this underestimation is the MMU spatial filter, which removed 29,801 km
2 (24.8%) of initially labeled secondary vegetation (
Table 4), part of which is true secondary vegetation below 2 ha that is systematically excluded from the product. Additions during post-processing and manual refinement (11,144 km
2;
Table 5) partially compensate this loss but do not close the gap between mapped and estimated area.
To position the present results within the operational product line,
Table 7 compares the 2024 secondary vegetation and pasture accuracies with the official assessments of the previous, manually produced TerraClass Cerrado editions [
62]. For secondary vegetation, the 2024 semi-automated product reaches a user’s accuracy of 96.27%, above the 78–88% range of the 2018–2022 editions, and a producer’s accuracy of 79.22%, above the 49–74% range of those editions; the pasture class remains comparably high. These editions used independent validation samples, reference years, and fully manual interpretation, so this is a cross-cycle comparison of official products rather than a controlled experiment. Within that limitation, the comparison indicates that the proposed workflow maintained program-level thematic accuracy while improving the secondary-vegetation result and replacing full manual delineation with cluster-level labeling.
This asymmetry between user’s and producer’s accuracy represents a design choice consistent with the requirements of operational monitoring programs. In TerraClass and PRODES, false positives directly affect land management decisions and policy enforcement, while omitted small fragments have lower operational impact. The high user’s accuracy indicates that mapped secondary vegetation areas are reliable for downstream applications, while the lower producer’s accuracy reflects the conservative approach applied across the entire Cerrado. The following section traces each of the 25 final errors through the processing stages to identify their specific sources.
4.4. Omission and Commission Trajectory
The final product contains 25 errors among the 695 validation points: 16 omissions and 9 commissions. Tracing each error through the processing trajectory reveals its specific origin, allowing a distinction between errors attributable to pipeline constraints and those reflecting the inherent limits of the SOM-HCA approach. These proportions are exploratory and specific to this validation sample; they indicate where errors occurred in the pipeline and should not be interpreted as general error rates for the method.
Cross-referencing the 16 omission errors with the processing trajectory (
Table 8) reveals two distinct patterns: 9 points (56.3%) were never detected as secondary vegetation at any stage, while the remaining 7 were detected but subsequently removed during post-processing (spatial filter: 3, manual review: 3, gap-fill reversion: 1).
The 9 omission points not detected as secondary vegetation at the center pixel were analyzed using a 5 × 5 pixel neighborhood (∼50 m) in the original cluster map (
Table 9). In 7 of these 9 cases, the neighborhood contained at least one secondary vegetation pixel, suggesting transition areas or fragments smaller than 2 ha.
Separating these two types of omission is important: those caused by post-processing decisions (spatial filter or manual removal of secondary vegetation initially estimated as present) and those where secondary vegetation was never estimated at any stage. The first type is inherent to any pipeline that applies spatial filtering constraints. The second type, limited to two cases, represents the actual estimation limit of the SOM-HCA approach for the temporal windows and spectral features used.
Figure 7 and
Figure 8 present complementary views of the 9 points never detected as secondary vegetation.
Figure 7 shows the SOM-HCA cluster mosaic before thematic labeling, allowing visual inspection of the spatial arrangement of spectro-temporal clusters around each validation point.
Figure 8 shows the same windows after the full processing trajectory, distinguishing pixels retained as secondary vegetation, pixels removed by the spatial filter or manual review, pixels added during refinement, and pasture/background. This comparison helps identify whether an omission reflects absence of a secondary-vegetation cluster near the point or a subsequent post-processing or removal decision. In panel (d) of
Figure 8, for example, the surrounding secondary-vegetation patch was identified during interpretation, but the validation point remains in the final pasture/background class after post-processing, indicating a possible true secondary-vegetation fragment removed by the refinement pipeline. Seven of these points are located at the boundaries of secondary vegetation patches, with secondary vegetation detected in neighboring pixels but not at the validation point. The remaining 2 had no secondary vegetation pixels within 50 m.
Commission errors are concentrated in the concordance trajectory (
Table 10), meaning that most false positives originated at the clustering stage and persisted through all subsequent stages without being corrected. Of the 9 commission errors, 5 (55.6%) originated from cluster labeling and persisted through all stages. In addition, 2 (22.2%) were introduced by gap filling, 1 (11.1%) by manual addition, and 1 (11.1%) by post-MMU recovery.
These cases correspond to pasture pixels whose spectro-temporal signatures are close enough to secondary vegetation patterns to be grouped into the same clusters. This type of confusion has been documented in previous Cerrado mapping efforts [
63] and remains a persistent challenge even in high-resolution multi-temporal analyses [
64]. It reflects the ecological continuum between regenerating vegetation and managed pastures, where spectral boundaries are gradual rather than discrete. The low number of commissions introduced by post-processing or manual editing indicates that the refinement stages do not amplify this confusion.
Overall, of the 25 final errors, 14 of 16 omissions (87.5%) had secondary vegetation detected by the SOM-HCA clustering either at the center pixel or in the immediate neighborhood, and only 2 (12.5%) were never estimated as secondary vegetation at any stage. Of the 9 commissions, 5 originated from cluster labeling and 4 were introduced during post-processing or manual editing.
4.5. Scalability and Cluster Labeling Strategy
Mapping secondary vegetation across the Cerrado requires processing approximately 692,000 km
2 of deforested land distributed across 20 ecoregions with distinct vegetation physiognomies and phenological patterns. At Sentinel-2 resolution (10 m), this represents billions of pixels per temporal composite. Supervised classification approaches face two scaling constraints in this context: the computational cost of training classifiers on datasets large enough to represent the biome’s heterogeneity, and the labor cost of collecting training samples across all ecoregions and temporal windows [
19,
20]. The pipeline presented here addresses both constraints through two mechanisms: WTA-based sampling that reduces the SOM training input to approximately 1% of the pixel population, and cluster-level labeling that replaces per-pixel sample collection with a fixed set of interpretation decisions.
The workflow converts a large spectro-temporal mapping problem into a tractable interpretation task through successive reductions of complexity: ecoregion and mask constraints narrow the domain, WTA-based sampling reduces the pixel population, SOM converts full trajectories into prototypes, and HCA-DTW groups these prototypes into interpretable clusters. The binary output class does not imply a binary feature space; it is the final operational label assigned after organizing a continuous secondary vegetation–pasture gradient into interpretable spectro-temporal patterns.
Supervised methods require 10,000 to 50,000 labeled samples for Cerrado-scale mapping [
19,
20], and classification quality depends on how well samples represent the spatial and temporal variability of target classes [
22]. This dependence is not limited to sample size, since training-data quality, sampling design, class balance, and class distribution are also recognized sources of variation in land-cover classification performance [
65]. In the Cerrado, gradual spectral transitions between vegetation physiognomies limit sample transferability across ecoregions and temporal windows [
36], requiring new collection for each mapping cycle. Supervised workflows also involve iterative cycles of training, evaluation, and sample refinement [
21], plus post-classification adjustments that are rarely quantified [
64] but operationally comparable to the refinement stage described here.
Published supervised and deep learning studies provide relevant context for interpreting the proposed workflow, although their direct comparison with TerraClass Phase 3 is limited by differences in target definition, reference data, study area, legend structure, and validation protocol. Existing Cerrado studies address different targets, such as deforestation detection with deep learning [
66], native/non-native vegetation and physiognomy mapping in sub-biome areas [
67], or natural-vegetation mapping with synthetic aperture radar (SAR)–optical deep learning [
68]. Their results show that performance depends strongly on label granularity, reference data, study area, and vegetation structure, with accuracy decreasing when Cerrado vegetation is disaggregated into physiognomic or structurally similar classes. Deep-learning physiognomy mapping in the Cerrado has so far been demonstrated mainly in smaller or curated study settings [
69], rather than as biome-scale operational products. These studies indicate that the secondary vegetation–pasture gradient is a difficult separability problem within the broader context of Cerrado vegetation mapping.
The clustering approach replaces sample collection and iterative refinement with approximately 3000 labeling decisions (150 clusters × 20 ecoregions), a fixed number independent of pixel count or area covered. The 150 clusters per ecoregion were determined empirically and applied uniformly. This standardization enables parallel processing with consistent protocols, but implies that some ecoregions may be over-partitioned while others are under-partitioned relative to their spectral complexity, as reflected in the concordance variability across ecoregions (
Supplementary Materials, Table S1).
The labeling process exhibits a learning curve: interpretations improve as specialists accumulate experience across ecoregions [
70], and inter-interpreter variability is managed through a quality control protocol with two independent analysts and a senior auditor [
71]. Future iterations could incorporate quantitative separation metrics to identify, before labeling, which ecoregions require more manual correction.
The operational impact of this approach is illustrated by the TerraClass Cerrado 2024 cycle. In previous cycles, Phase 3 (secondary vegetation) relied entirely on pixel-level visual interpretation, requiring analysts to manually delineate secondary vegetation polygons across the entire 692,000 km
2 study area. With the proposed methodology, the clustering stage provides an initial map, with 86.3% of the final product originating from the initially labeled clusters (
Table 5), shifting the analyst’s role from full-coverage polygon creation to targeted review and correction of a pre-classified product.
4.6. MMU Spatial Filter Impact
The 2-hectare MMU is the single largest source of area reduction in the pipeline (
Table 4). This filter is not specific to the proposed methodology; it is an operational requirement for interoperability with PRODES and TerraClass [
6].
The clustering approach, however, may be more susceptible to MMU-based losses than methods that produce smoother spatial outputs. Because SOM-HCA assigns each pixel to a cluster based on spectro-temporal similarity, pixels along vegetation transitions are often grouped into different clusters than adjacent secondary vegetation pixels. When the transition cluster is not labeled as secondary vegetation, the resulting gap fragments the secondary vegetation patch into smaller pieces that may fall below 2 ha. The spatial filter removed 29,801 km
2 (24.8%) of the initially labeled secondary vegetation area (
Table 4), and part of this removal corresponds to true secondary vegetation that falls below 2 ha. The post-MMU recovery mechanism partially compensates this loss: interpreters reinstated 1656 km
2 (2.0% of the final secondary vegetation product;
Table 5) of fragments judged to be part of larger secondary vegetation patches, but this represents only a fraction of the area excluded by the spatial filter.
The omission error trajectory (
Table 8) confirms that the MMU constraint contributes directly to the final omission count: 3 of the 16 omission errors (18.8%) correspond to secondary vegetation points correctly identified by the clustering but removed by the spatial filter. Combined with the 3 points removed manually and 1 reverted after gap filling, 7 of the 16 omissions (43.8%) represent areas where secondary vegetation was initially estimated as present and subsequently excluded during refinement. These are not estimation failures of the SOM-HCA approach; they are consequences of the post-processing constraints applied to the classification output.
4.7. Limitations
The selection of temporal windows represents an operational design choice that may affect subsequent workflow steps. The four temporal windows used across ecoregions were selected to approximate dry-season conditions while meeting the TerraClass production schedule. Although the ecoregion-specific assignment was designed to minimize temporal-window effects, different windows can expose different phenological states and therefore alter the spectro-temporal patterns available to SOM-HCA clustering. This may influence cluster composition within each ecoregion, the initial cluster labels, and the amount of post-processing or manual refinement required.
These workflow reductions also introduce limitations and methodological trade-offs. The workflow was validated as an integrated final product; the isolated contribution of each reduction step was not evaluated separately. Ecoregion-based processing reduces the number of spectro-temporal patterns represented within each SOM-HCA run and makes the analysis computationally feasible, but the independently produced ecoregion outputs must later be assembled into a biome-scale product. These ecoregion-level differences may affect spatial consistency along the boundaries between ecoregion-level outputs. This limitation is partly reduced by the window-selection strategy described in
Section 2.3.1, which considered regional climatic dynamics, the availability of cloud-free images, dominant phytophysiognomies, and their seasonal behavior, and by the regionalization strategy described in
Section 2.3.3, in which ecoregions were treated as broad processing units connected by continuous ecological and environmental gradients.
In this workflow, masking, regionalization, and spatial filtering are connected rather than independent, and their effects can interact with temporal-window selection. Masking reduces complexity by restricting clustering to the TerraClass Phase 3 domain, avoiding unrelated classes and reducing the number of spectro-temporal patterns assigned to SOM-HCA. However, it also constrains the spatial domain before clustering, so boundary effects from the mask can interact with the later MMU filter. The MMU, in turn, removes noise and enforces the operational 2-ha mapping standard, but it may also exclude true secondary-vegetation fragments below the threshold. As a result, fragmentation in the initial map, MMU removals, and subsequent manual corrections may arise from the combined effects of temporal-window selection, masking, regionalization, and spatial filtering rather than from any single processing step.
The WTA temporal reduction is a reduced scalar projection used only to organize sampling. Time series with close time-weighted means but different temporal shapes or phenological phases may fall in the same bin. This limitation is partly controlled by the stratified sampling design, in which bins are defined from the empirical WTA distribution, samples are allocated according to stratum frequency, and pixels are selected randomly within each stratum. The complete SITS information is preserved for SOM training, HCA, and DTW, so this limitation is restricted to sample selection; however, the sampling stage cannot guarantee representation of every temporal behavior.
The sampling strata were defined by equal-width binning of the WTA range, which is computationally simple but does not adapt to the empirical distribution of WTA values. Future work should evaluate whether distribution-aware discretization strategies can better account for this distribution, particularly for the representation of sparsely populated strata and minority temporal patterns.
The physiognomic and spectro-temporal complexity of Cerrado vegetation makes fully homogeneous clusters difficult to obtain, as distinct vegetation types may exhibit similar temporal trajectories and overlap with pasture [
36]. In the current protocol, cluster interpretation addresses this ambiguity through a configuration calibrated from a complex pilot ecoregion and an inclusive selection of clusters with a substantial secondary-vegetation component, even when some mixture is present. The variability in concordance and refinement effort across ecoregions also suggests a limitation of applying a standardized parameter configuration to heterogeneous landscapes. Although the SOM grid and cluster count were calibrated in a high-complexity pilot ecoregion, this calibration did not fully account for differences in local spectro-temporal complexity across the Cerrado. Future work should therefore evaluate ecoregion-adaptive parameterization, including WTA subdivision, SOM grid size, sampling fraction, and cluster count, to improve local class separability and reduce refinement effort.
The methodology relies on external masks from PRODES and TerraClass to restrict the analysis domain. Temporal mismatches between the mask reference year and the mapping period introduce classification errors in ecoregions with active agricultural frontiers. Additionally, although the validity mask incorporates the previous year’s agriculture layer, the current-year agriculture mask was produced concurrently with secondary vegetation mapping and therefore could not be used as an a priori filter during cluster labeling. This limitation resulted in redundant editing concentrated in frontier ecoregions (e.g., Basaltos do Paraná at ∼26%;
Supplementary Materials, Table S3). Applying the agriculture mask before cluster labeling would reduce this redundancy but would require producing the agriculture layer first, a phase dependency not feasible in the current cycle.
Finally, the methodology operates in a binary classification context (secondary vegetation vs. pasture) and does not distinguish successional stages or vegetation structure within secondary vegetation. The SOM-HCA clusters encode spectro-temporal variability within secondary vegetation, but this information is collapsed into a single label. Extending the labeling protocol to subclasses would increase the product’s information content but also increase interpretation complexity.
4.8. Operational Implications for Monitoring Programs
The ecoregion-based strategy enables independent, parallel processing: each ecoregion enters the pipeline as imagery becomes available, multiple teams work concurrently, and reprocessing one ecoregion does not affect the others. This compresses the mapping cycle from a sequential effort spanning years into a coordinated campaign of months.
The clustering stage provides the foundation for this acceleration. Rather than constructing the vegetation map from scratch through pixel-level visual interpretation, analysts receive a pre-classified baseline in which 86.3% of the final secondary vegetation area traces back to the original cluster labeling (
Table 5), with post-processing and manual editing contributing 5.1% and 6.6%, respectively. The analyst’s task shifts from full-coverage polygon delineation to targeted review of boundaries, ambiguous transitions, and areas the algorithm missed. Because the clustering is calibrated to overestimate secondary vegetation extent, borderline areas are flagged for expert review rather than silently omitted, since correcting omissions in later stages is more costly than removing commissions during editing.
The domain restrictions imposed by the initial masking further improve efficiency. By excluding primary vegetation, agriculture, urban areas, and water bodies, the clustering operates on a narrower attribute space and achieves better class separability. Successive mapping cycles reinforce this effect: as previously mapped polygons are incorporated into the validity mask, the analysis domain contracts around the active regeneration frontier, progressively narrowing the classification problem.
The TerraClass Cerrado 2024 cycle illustrates the practical impact. Phase 3 (secondary vegetation) was completed between June and November 2024, whereas previous cycles based entirely on manual interpretation typically required two years. This shorter path from satellite acquisition to finished product enables more frequent updates of land use trajectories combining deforestation (PRODES), agriculture (TerraClass), and secondary vegetation, allowing decision-makers to detect regeneration fronts, evaluate restoration progress, and identify renewed conversion while these processes are still underway. This responsiveness directly supports monitoring of Brazil’s Nationally Determined Contributions under the Paris Agreement and the restoration targets of the National Plan for Native Vegetation Recovery (Planaveg) [
3,
12].
Although the present product targets the secondary vegetation and pasture classes of TerraClass Phase 3, the workflow can support other land-use and land-cover targets defined by temporal behavior, such as deforestation, cropland expansion, and seasonally flooded or wetland areas, after adapting the labeling protocol to the class of interest.
5. Conclusions
This study presented a scalable methodology for secondary vegetation mapping that combines unsupervised clustering with expert interpretation across large and heterogeneous areas. Cluster-level labeling limited the initial interpretation task to approximately 3000 cluster-level decisions across 20 ecoregions. Applied to 692,000 km2 of previously deforested land in the Brazilian Cerrado, the methodology produced a secondary vegetation map of 81,209 km2 (11.74%), with 95% confidence intervals of 98,683 ± 10,071 km2 for estimated area, 96.45 ± 1.52% for overall accuracy, 96.27 ± 2.40% for secondary vegetation user’s accuracy, and 79.22 ± 7.94% for secondary vegetation producer’s accuracy, corresponding to an F1-score of 86.90%.
The SOM-HCA map obtained after cluster labeling accounted for 86.3% of the final secondary vegetation area. Taken as an operational proxy for editing intensity, this final-map composition indicates that most mapped secondary vegetation was derived directly from initial cluster labeling, while manual refinement concentrated on excluding 17.0% of the initially labeled area and adding 6.6% of the final product. This shifted the TerraClass Phase 3 workflow from full manual delineation to targeted refinement of a pre-classified map, helping explain the observed reduction in production time. Future work should evaluate ecoregion-adaptive parameterization, including grid size, cluster count, sampling design, spectral feature selection, and complementary data sources according to local landscape complexity.
Phase 3 (secondary vegetation) of the TerraClass Cerrado 2024 cycle was completed in six months (June–November 2024), whereas previous cycles based on manual polygon delineation required about two years. The operational gain extended beyond processing time. For the secondary vegetation class, the 2024 product reached a user’s accuracy of 96.27% and a producer’s accuracy of 79.22%, above the user’s accuracy (78 to 88%) and producer’s accuracy (49 to 74%) of the manually produced 2018, 2020, and 2022 editions (
Table 7). The workflow therefore shortened the cycle while preserving program-level thematic accuracy and improving it for the secondary vegetation class. Within the 2024 cycle, the approach was operationally feasible for biome-scale secondary vegetation mapping, and with further evaluation across additional cycles it may support more frequent land-use updates for tracking secondary vegetation dynamics, evaluating restoration progress, and informing monitoring of national commitments such as the Nationally Determined Contributions under the Paris Agreement and the National Plan for Native Vegetation Recovery (Planaveg).