3.2. Airport State Identification Results
To identify stage-specific differences in the monthly operational states of major airports in Japan, this study implemented the clustering analysis using Python’s scikit-learn library. The raw route-level operational data were first aggregated to the airport-month level, after which the eight state-identification indicators defined in Equations (1)–(8) were calculated. The 2010 observations were retained as the previous-year baseline for calculating the year-on-year indicators for 2011. After excluding observations without a valid previous-year comparison, the final clustering dataset contained 1113 airport-month observations covering January 2011 to March 2024. All indicators were standardized according to Equation (9), and K-means clustering was subsequently applied to the standardized feature set. Candidate solutions from K = 2 to K = 8 were evaluated using four complementary internal validation criteria: the elbow criterion, the average silhouette score, the Calinski–Harabasz index, and the Davies–Bouldin index. Based on the joint consideration of quantitative clustering validity and substantive interpretability, K = 3 was selected as the preferred clustering solution. The cluster distribution, mean feature values, PCA scatter plot, and radar chart were then used to support the substantive interpretation of the identified operational states.
3.2.1. Determination of the Number of Clusters and Clustering Validation
Table 6 reports the internal validity metrics for candidate clustering solutions from K = 2 to K = 8. As K increases, WCSS decreases continuously, while the percentage reduction in WCSS becomes progressively smaller, indicating diminishing returns in within-cluster compactness. The average silhouette score reaches its highest value at K = 3 (0.4886), whereas the Calinski–Harabasz index is slightly higher at K = 2 but remains comparably high at K = 3. The Davies–Bouldin index is lower for some higher-K solutions, but these solutions exhibit weaker separation and poorer substantive interpretability. Considering the elbow trend, the silhouette score, the Calinski–Harabasz index, the Davies–Bouldin index, and the interpretability of the resulting states, K = 3 was selected as the final clustering solution.
As shown in
Table 6 and
Figure 1, K = 3 provides the most balanced solution overall in terms of cluster compactness, separation, and substantive interpretability. Although K = 2 yields a slightly higher Calinski–Harabasz index and K = 6 produces a slightly lower Davies–Bouldin index, K = 3 achieves the highest average silhouette score among all candidate solutions and avoids the marked deterioration in separability observed for K = 4 and higher-K solutions. Therefore, considering both quantitative validity and operational interpretability, K = 3 was selected as the final clustering solution.
Compared with adjacent candidate solutions, the K = 2 solution under-clustered the data by merging substantively distinct airport operational conditions, whereas the K = 4 solution over-clustered the data by splitting the recovery state into highly similar subgroups with limited additional interpretive value. Therefore, the K = 3 solution offers the best compromise between statistical validity and substantive interpretability. A detailed comparison of the clustering structures for K = 2, K = 3, and K = 4 is provided in
Appendix A.1.
3.2.2. Clustering Results and State Classification
The monthly state clustering results for all airports are summarized in
Table 7. The results show substantial differences in the number of airport-month observations assigned to each state, indicating that airport operational states are not uniformly distributed over the sample period but exhibit a clear stage-specific structure.
The K = 3 solution classified the 1113 airport-month observations into 175 low-activity-state observations, 852 recovery-state observations, and 86 rapid-rebound-state observations, accounting for 15.72%, 76.55%, and 7.73% of the full clustering sample, respectively. The recovery state therefore represents the predominant operational pattern in the sample, whereas rapid-rebound episodes occur relatively infrequently.
From the perspective of sample distribution, the recovery state is overwhelmingly dominant, suggesting that, for most months during the study period, major Japanese airports remained in a relatively stable operational regime. By contrast, the low-activity and rapid-rebound states account for much smaller proportions, indicating that severe operational contraction and unusually strong rebound were concentrated in specific periods rather than representing normal operating conditions. Overall, the recovery state constitutes the predominant and relatively stable operational condition, the low-activity state primarily corresponds to periods of depressed operations under external shocks, and the rapid-rebound state represents relatively infrequent episodes of exceptionally strong year-on-year growth from a comparatively low operational base.
3.2.3. Differences in State Characteristics and Their Economic Interpretation
To further interpret the substantive meaning of the K-means clustering results, this study compares the mean values of the three identified states across eight indicators, including year-on-year passenger growth, year-on-year flight growth, year-on-year cargo growth, seat utilization rate, weight utilization rate, and the 2019-relative levels of passenger traffic, flights, and cargo. The results are reported in
Table 8.
The cluster-specific mean values reported in
Table 8 reveal distinct operational profiles across the three identified states. The low-activity state is characterized by simultaneous contractions in passenger traffic, flight movements, and cargo throughput, with mean year-on-year growth rates of −29.70%, −15.34%, and −25.72%, respectively. Its seat and weight utilization rates are also relatively low, at 49.65% and 35.41%, respectively. In addition, its passenger, flight, and cargo levels correspond to only 41.59%, 64.43%, and 63.97% of their respective 2019 benchmarks, indicating a substantial reduction in overall airport activity.
The recovery state exhibits moderate positive growth in passenger traffic, flight movements, and cargo throughput, with mean year-on-year growth rates of 9.51%, 4.15%, and 2.77%, respectively. It records the highest seat and weight utilization rates among the three states, at 74.40% and 50.26%, respectively. Its passenger and flight levels reach 93.31% and 99.42% of their corresponding 2019 levels, while its cargo level reaches 105.80%. These characteristics indicate relatively stable airport operations that have returned close to the pre-shock benchmark or, in the case of cargo, slightly exceeded it.
The rapid-rebound state is distinguished primarily by exceptionally strong year-on-year growth, with passenger, flight, and cargo growth rates reaching 135.61%, 72.29%, and 87.34%, respectively. However, its seat and weight utilization rates, at 61.54% and 43.29%, remain below those of the recovery state. Its passenger, flight, and cargo levels are equivalent to 64.97%, 87.71%, and 62.00% of their corresponding 2019 levels, respectively. Therefore, the rapid-rebound state reflects a rapid rebound or expansion from a comparatively low operational base, rather than the highest absolute level of airport activity.
To further examine and validate the plausibility of the clustering results, this study supplements the comparison of mean feature values with a PCA scatter plot and a radar chart of state-specific feature means, thereby providing additional evidence on the separability of the three airport operational states and the structural differences in their indicator profiles. The results are shown in
Figure 2.
The PCA scatter plot indicates that the three states exhibit substantial, although not complete, separation in the principal-component space. The low-activity and rapid-rebound states occupy comparatively distinct regions, whereas recovery-state observations are concentrated mainly in the central region and partially overlap with the other two states. This visual pattern is consistent with the internal validation results, including the average silhouette score of 0.4886 for K = 3, which indicates meaningful but not perfect cluster separation. The radar chart further shows that the low-activity state is characterized by negative or weak growth, low utilization rates, and low 2019-relative operational levels; the recovery state records the highest utilization rates and operational levels closest to the 2019 benchmark; and the rapid-rebound state is distinguished primarily by exceptionally strong year-on-year growth from a comparatively depressed operational base. Taken together, the two visualizations support the substantive interpretation of the three clusters, while also indicating that the identified operational states should be understood as empirically distinguishable but not completely discrete operational regimes.
3.2.4. Temporal Evolution Characteristics of Airport Operational States
After identifying the three airport operational states, the next step is to examine their evolution over time. To this end, this study analyzes the stage-specific temporal dynamics of major Japanese airports from two complementary perspectives: the state trajectories of individual airports and the overall distribution of states across airports over time. Specifically, the analysis is based on airport-specific state sequence plots and a heatmap of airport states. The corresponding airport-specific state sequences and the overall state heatmap are presented in
Figure 3 and
Figure 4, respectively.
The airport-specific state sequence plots reveal pronounced stage-specific evolution in the operational states of major Japanese airports. During 2011–2019, most airports remained predominantly in the recovery state for the majority of months, with only occasional fluctuations into the low-activity or rapid-rebound states. This suggests that, prior to the pandemic, airport operations in Japan were generally stable, with the recovery state constituting the dominant long-term operating condition.
From 2020 onward, however, airport operational states changed markedly, as all seven airports shifted into the low-activity state within a relatively similar time frame, indicating that the airport system was substantially affected by external shocks. After the shock, airport operations did not immediately return to a stable condition; instead, most observations gradually shifted from the low-activity state to the recovery state, while some months were classified as belonging to the rapid-rebound state because they exhibited exceptionally strong year-on-year growth and rapid rebound from a lower operational base. The repeated occurrence of the rapid-rebound state around 2021–2022 therefore reflects episodic periods of rapid rebound or expansion rather than the attainment of the highest absolute operational level. By 2023, most airports had returned predominantly to the recovery state, indicating that overall airport operations were gradually regaining stability.
The inclusion of the COVID-19 period is important for identifying the full range of airport operational states, especially the low-activity state and the subsequent recovery process. Nevertheless, the pandemic represents an exceptional and system-wide external shock rather than a routine source of operational fluctuation. Therefore, the state evolution patterns observed during this period should not be interpreted as fully representative of normal operating conditions. In particular, the estimated classification performance may be partly shaped by the sharp contraction and subsequent recovery dynamics associated with the pandemic period. Accordingly, these pandemic-related dynamics are taken into account when interpreting the classification results and are further acknowledged as a limitation of the present study. A sensitivity analysis excluding the pandemic and immediate recovery periods was not conducted because such exclusion would remove most low-activity and rapid-rebound observations, making the resulting state structure and minority-class evaluation statistically unstable.
The results show that, following the external shock, all airports synchronously entered the low-activity state for a prolonged period, exhibiting a high degree of overall consistency. This suggests that the shock had a broad and pervasive impact on the major airport system in Japan. At the same time, airports differed in the timing of their return to the recovery state and in the frequency with which they experienced rapid-rebound episodes, reflecting inter-airport differences in recovery rhythm, fluctuation intensity, and the speed of state transition. Among them, Narita and Kansai airports exhibited the rapid-rebound state relatively more frequently, indicating more pronounced stage-specific rebound characteristics, whereas Haneda, Fukuoka, and Osaka airports returned earlier to an operational pattern dominated by the recovery state, demonstrating relatively stronger stability.
3.2.5. Transition Patterns of Airport Operational States
To further reveal the dynamic transition relationships among airport operational states, this study constructs a state transition matrix and visualizes it using a heatmap, as shown in
Figure 5, to analyze the transition probabilities among the low-activity, recovery, and rapid-rebound states.
The state transition matrix shows that all three states exhibit some degree of persistence, although their stability differs markedly. The recovery state has the highest self-transition probability, reaching 95.50%, and therefore represents the most persistent operational regime. The self-transition probability of the low-activity state is 82.86%, suggesting that once airports enter a severely depressed operational condition, they tend to remain there for a certain period rather than recover immediately. By contrast, the rapid-rebound state has a substantially lower self-transition probability of 67.44%, confirming that it more often represents a temporary high-growth episode from a comparatively low base than a long-term stable operating condition.
In terms of transition direction, the probability of moving from the low-activity state to the recovery state is 12.00%, which is substantially higher than the probability of moving directly from the low-activity state to the rapid-rebound state (5.14%). This result indicates that observations emerging from severe operational contraction are more likely to enter the relatively stable recovery state than to shift directly into a rapid-rebound episode. The probability of moving from the rapid-rebound state to the recovery state is 24.42%, which is considerably higher than the probability of moving from the rapid-rebound state to the low-activity state (8.14%). Rapid-rebound episodes are therefore more likely to return to the stable recovery regime after a period of unusually strong growth than to deteriorate directly into severe contraction. Overall, the transition structure is centered on a highly persistent recovery state, whereas rapid-rebound episodes are less frequent, less stable, and often followed by a return to recovery. The three states should therefore not be interpreted as a fixed ordinal or linear progression from low activity to recovery and then to a higher operational state.
3.3. Evaluation and Analysis of the Retrospective Airport-State Classification Models
Building on the full-sample state-identification and transition-analysis results, the supervised component evaluates whether retrospectively identified recovery and rapid-rebound labels can be distinguished using lagged operational features. The 60/20/20 chronological partition is applied only to the supervised-classifier stage. Because the target labels were generated through pooled standardization and K-means clustering fitted to the complete sample, and because the 2019-relative variables for pre-2019 observations were constructed using ex post information, the reported metrics assess how well the supervised models reproduce retrospectively identified full-sample labels in chronologically later observations. They do not constitute end-to-end out-of-sample validation or prospective forecasting accuracy.
3.3.1. Comparison of the Overall Classification Performance of Different Models
To systematically compare the classification performance of different learning mechanisms, six classification models were evaluated on the same classifier-stage test subset using accuracy, macro-precision, macro-recall, and macro-F1 score. All models were trained and evaluated under the same chronological data partitioning and class-imbalance treatment procedure. The comparative results are presented in
Figure 6.
As shown in
Figure 6, the six models exhibit considerable differences across the four evaluation metrics. Gradient boosting achieves the highest overall accuracy of 0.8485, followed by logistic regression at 0.8364 and random forest at 0.8242. However, accuracy alone is insufficient for evaluating the present task because the classifier-stage test subset remains imbalanced and the rapid-rebound state represents the minority class.
In terms of balanced classification performance on the current classifier-stage test subset, LightGBM achieves the highest macro-recall of 0.6986 and the highest macro-F1 score of 0.6864 among the six evaluated models. Although its overall accuracy of 0.7818 is lower than that of gradient boosting, logistic regression, and random forest, its higher macro-recall and macro-F1 indicate a more balanced ability to identify both operational states. Gradient boosting ranks second in macro-F1, with a value of 0.6778, whereas XGBoost obtains a macro-F1 score of 0.6387.
Logistic regression and random forest achieve relatively high macro-precision values of 0.9147 and 0.9088, respectively, but their lower macro-recall values indicate that these models produce relatively conservative predictions and fail to identify a substantial proportion of minority-class observations. The RBF-SVM model performs least effectively, particularly in terms of macro-precision and macro-F1. Overall, within the primary SMOTE-based six-model comparison, the evaluated boosting-based models yielded higher point estimates for balanced classification performance than the evaluated linear, kernel-based, and bagging-based models.
To further assess the uncertainty surrounding these point estimates, stratified bootstrap confidence intervals were calculated for the main classification metrics using predictions from the chronologically later test subset. The corresponding results are reported in
Table 9. Compared with point estimates alone, these confidence intervals provide additional evidence regarding the stability and uncertainty of the observed performance differences among the six models.
Values are reported as point estimates [95% bootstrap confidence intervals] for the primary SMOTE-based six-model comparison. The confidence intervals were obtained using 1000 stratified bootstrap resamples of the predictions on the chronologically later test subset. Because individual airport-month observations were resampled within each class, the procedure did not explicitly preserve serial dependence within airports or common month-level dependence across airports. The intervals should therefore be interpreted as conditional empirical uncertainty estimates under the adopted resampling scheme. The degenerate confidence intervals for RBF-SVM arise because the model predicted all test observations as the recovery state, resulting in invariant bootstrap metrics under the stratified resampling procedure.
As shown in
Table 9, Gradient boosting obtained the highest point estimate for overall accuracy, whereas LightGBM obtained the highest point estimates for macro-recall, macro-F1, and recall for the rapid-rebound state. However, the confidence intervals indicate that some differences between models, particularly those with similar point estimates, should be interpreted cautiously. The model comparison is therefore reported as evidence of comparative classification performance on the current classifier-stage test subset rather than as definitive proof of statistically universal superiority. The paired bootstrap comparisons between LightGBM and the benchmark models are further reported in
Appendix A.3,
Table A3.
The reported classification performance should be interpreted as classifier-stage performance on chronologically later observations within the present retrospectively constructed dataset. It does not demonstrate that equivalent performance would be achieved in a prospective real-time application using only information available at each historical decision date.
Because the classification objective is not only to maximize overall accuracy but also to identify observations belonging to the less frequent rapid-rebound state, the class-specific recognition performance of the six models is examined further in the following subsection.
3.3.2. Analysis of Model Capability in Identifying the Rapid-Rebound State
Because rapid-rebound states represent relatively infrequent episodes of exceptionally strong year-on-year growth from a comparatively low operational base, identifying them is useful for distinguishing short-term rebound intensity from stable near-normal operation. Accordingly, this subsection focuses on class-specific recognition performance rather than overall accuracy alone. The results show that the six algorithms differ considerably in their ability to identify the minority rapid-rebound state.
Although several models achieve relatively high recovery-state recognition accuracy because of the larger number of recovery observations, their ability to identify rapid-rebound observations varies substantially, as reflected by the class-specific recall and F1-score comparisons shown in
Figure 7.
The zero recall and F1-score values observed for some models indicate that none of the rapid-rebound observations were correctly identified in the classifier-stage test subset. Under the standard definitions of these metrics, recall becomes zero when the number of true positives is zero, and the corresponding F1-score is consequently also zero. These values therefore reflect model classification behavior rather than a metric-calculation failure. Although SMOTE balances the class counts in the training subset through synthetic oversampling, the original training subset contains only seven rapid-rebound observations, which limits the diversity of minority-state patterns available for model learning. Moreover, SMOTE-generated samples cannot provide the same independent information as additional real observations, while the validation and test subsets retain their original imbalanced distributions. Consequently, the zero values reveal the tendency of some classifiers to favor the dominant recovery state and their limited ability to generalize to rare rapid-rebound observations.
The comparatively stronger performance observed for LightGBM in the current experiment may be related to its boosting-based tree structure, which is designed to iteratively focus on difficult-to-classify samples and model nonlinear interactions among lagged growth indicators, operational efficiency variables, and 2019-relative operational-level features. In the present dataset, logistic regression and random forest showed lower minority-class recognition performance than LightGBM. This difference may be associated with the linear specification of logistic regression and the distinct learning mechanism of the bagging-based random forest model.
These results are consistent with the possibility that the retrospective distinction between recovery and rapid-rebound observations involves nonlinear temporal relationships that are not fully captured by the evaluated linear model. Accordingly, models capable of representing nonlinear interactions may assist in the retrospective monitoring of rapid-rebound episodes under data conditions similar to those examined in this study.
3.3.3. Classification Performance and SHAP-Based Interpretation of the LightGBM Model
Based on the primary SMOTE-based six-model reference comparison, LightGBM was selected as the representative model for SHAP-based interpretation because it yielded the highest point estimates for macro-F1 and rapid-rebound-state recall among the six evaluated classifiers. Within this primary SMOTE-based comparison, the paired bootstrap results indicated that the 95% confidence intervals for LightGBM’s differences in rapid-rebound-state recall relative to the benchmark models excluded zero under the adopted stratified resampling scheme, whereas the confidence intervals for its macro-F1 differences relative to several models included zero. These results should be interpreted cautiously because the resampling procedure did not explicitly account for serial dependence within airports or common month-level dependence across airports. Confusion-matrix analysis and SHAP-based feature-attribution analysis were therefore conducted to characterize the classification behavior of the SMOTE-based LightGBM model and to examine how individual features contributed to its fitted distinction between the recovery and rapid-rebound state labels.
The row-normalized confusion matrix of the LightGBM model is presented in
Figure 8. It should be noted that the classifier-stage test subset retained its original class distribution and contained no SMOTE-generated observations. The 34 rapid-rebound observations were original airport-month records, and their relatively large number reflects the temporal concentration of rapid-rebound episodes in the chronologically later post-shock period. On the current classifier-stage test subset, the model correctly classified 110 of the 131 recovery-state observations, corresponding to a recovery-state recall of 0.8397, while 21 recovery-state observations were incorrectly classified as rapid rebound. Among the 34 rapid-rebound-state observations, 19 were correctly identified, corresponding to a rapid-rebound-state recall of 0.5588, whereas 15 were misclassified as recovery. Thus, the four cells in
Figure 8 correspond to 0.84 (
n = 110), 0.16 (
n = 21), 0.44 (
n = 15), and 0.56 (
n = 19), respectively. Because the classifier-stage test subset contains only 34 rapid-rebound-state observations, the rapid-rebound-state recall should be interpreted together with its 95% bootstrap confidence interval of [0.4118, 0.7059], rather than as an exact estimate of generalizable minority-state recognition performance.
The asymmetric classification pattern suggests that identifying rapid-rebound observations remains more challenging than identifying recovery observations. This result is plausible because the rapid-rebound state is defined primarily by unusually strong year-on-year growth from a comparatively low base, while some of its utilization rates and benchmark-relative operational levels may overlap with those observed during recovery. The partially shared feature distributions between the two states therefore make rapid-rebound observations more difficult to distinguish consistently.
Compared with the benchmark models that predominantly classified observations into the majority recovery state, LightGBM correctly identified a larger number of rapid-rebound observations in the current classifier-stage test subset. Within the present comparison, this result suggests that the fitted boosting-based model may represent interactions among lagged growth indicators, operational-efficiency variables, and 2019-relative operational-level features more effectively than the specific benchmark models evaluated in this study.
To further characterize how the fitted LightGBM model uses the input variables, SHAP analysis is employed to quantify the contribution of individual features to the fitted distinction between the recovery and rapid-rebound state labels. SHAP values reflect model-specific feature contributions and associations rather than causal effects. Accordingly, features with larger absolute SHAP values are interpreted as contributing more strongly to the fitted classification output, rather than as causally driving changes in airport operational states. The SHAP-based ranking of feature contributions is presented in
Figure 9.
The SHAP ranking indicates that short-term lagged growth indicators and twelve-month 2019-relative operational-level variables make relatively large contributions to the LightGBM classification results. In particular, the dominant contribution of the one-month lag of passenger growth is consistent with the defining characteristic of the rapid-rebound state: exceptionally strong year-on-year growth from a comparatively low operational base. The contributions of the twelve-month 2019-relative variables further indicate that the fitted model distinguishes rebound intensity jointly with the airport’s longer-term operational position relative to the 2019 benchmark. These SHAP results should therefore be interpreted as explaining the fitted distinction between recovery and rapid-rebound observations, rather than as indicating movement toward a higher absolute operational level.
3.3.4. Exploratory Sensitivity Analysis Under Alternative Class-Imbalance Handling Strategies
To examine the sensitivity of the classification point estimates to the selected class-imbalance treatment, an exploratory analysis was conducted using SMOTE, random oversampling, and class-weighted learning. Rather than repeating the complete six-model comparison, three representative classifiers were included to cover distinct learning mechanisms: logistic regression as a linear classifier, random forest as a bagging-based ensemble classifier, and LightGBM as a boosting-based classifier. For the random-oversampling and class-weighted variants, the hyperparameter configurations selected under the primary SMOTE-based validation procedure were retained, and no strategy-specific hyperparameter retuning was performed. All variants were evaluated on the same chronologically later test subset. Therefore, this analysis examines changes in point estimates under alternative imbalance treatments and should not be interpreted as a fully optimized comparison of imbalance-handling strategies or as a second model-selection procedure.
This exploratory sensitivity analysis was conducted after the primary SMOTE-based procedure had been fixed and was not used to redefine the primary model-selection process. Promoting the class-weighted specification to the primary analysis solely on the basis of its higher point estimates on the test subset would constitute post hoc outcome-driven model selection.
The exploratory sensitivity results are reported in
Table 10. For logistic regression, SMOTE and random oversampling produced identical point estimates, with an accuracy of 0.8364, a macro-F1 score of 0.6240, and rapid-rebound-state recall of 0.2059. Under class-weighted learning, macro-F1 decreased to 0.5539 and rapid-rebound-state recall decreased to 0.1176. Thus, within the fixed hyperparameter configuration used in this exploratory analysis, changing from resampling to inverse-frequency class weighting did not improve minority-state recognition for logistic regression. This result applies to the present implementation and should not be interpreted as general evidence against class weighting or as proof that the underlying airport-state distinction is inherently nonlinear.
Random forest produced identical point estimates under the three evaluated imbalance-handling strategies, with an accuracy of 0.8242, a macro-F1 score of 0.5784, and rapid-rebound-state recall of 0.1471. The model maintained relatively strong recognition of the recovery state but identified only a small proportion of rapid-rebound observations. Under the fixed hyperparameter configuration retained from the primary SMOTE-based procedure, the random forest results were therefore relatively insensitive to the three evaluated imbalance treatments, although minority-state recognition remained limited.
For LightGBM, SMOTE and random oversampling produced the same point estimates, including a macro-F1 score of 0.6864 and rapid-rebound-state recall of 0.5588. The class-weighted LightGBM specification yielded point estimates of 0.8970 for accuracy, 0.8609 for macro-F1, and 0.9412 for rapid-rebound-state recall. Its confusion matrix indicated that 32 of the 34 rapid-rebound observations were correctly classified, while recovery-state recall remained 0.8855. These point estimates were materially higher than those obtained under SMOTE and random oversampling. However, because the class-weighted specification retained hyperparameters selected under the primary SMOTE-based procedure and was not accompanied by strategy-specific retuning or equivalent bootstrap uncertainty assessment, the observed differences should not be interpreted as statistically established superiority of class-weighted LightGBM.
Overall, the exploratory analysis indicates that estimated classification performance, particularly rapid-rebound-state recall, is materially sensitive to the selected imbalance-handling strategy. The results do not confirm the robustness of the primary SMOTE-based estimates, nor do they establish class-weighted learning as the optimal strategy. Instead, they identify imbalance treatment as an important source of methodological uncertainty in the present setting, where only seven rapid-rebound observations were available in the original training subset. The higher point estimates obtained by class-weighted LightGBM suggest that direct loss weighting may warrant further investigation, but a definitive comparison would require strategy-specific hyperparameter tuning, application to the complete set of candidate classifiers, and equivalent confidence-interval or paired-bootstrap assessment.
From a managerial perspective, the rapid-rebound state should not be interpreted automatically as a superior operational condition or as evidence of complete recovery. Its exceptionally high year-on-year growth occurs alongside utilization rates and 2019-relative operational levels that remain below those of the recovery state. Airport managers should therefore distinguish rebound speed from recovery level and should jointly consider growth momentum, resource utilization, and benchmark-relative operational position when assessing recovery progress. A rapid-rebound classification may indicate a need to monitor whether short-term demand growth can be converted into a stable recovery regime; by itself, however, it does not justify permanent capacity expansion or long-term resource commitments.