Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (20,713)

Search Parameters:
Keywords = Random forest

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
20 pages, 1997 KB  
Article
Algorithmic Diffusion on YouTube: A Machine Learning Analysis of Channel-Level Information Spread and Its Cross-Platform Generalisability
by Dana Tyulemissova, Aigul Shaikhanova, Oleksandr Kuznetsov, Aigerim Sambetova, Kainizhamal Iklassova and Aisanim Sarsenbayeva
Mach. Learn. Knowl. Extr. 2026, 8(9), 272; https://doi.org/10.3390/make8090272 (registering DOI) - 6 Sep 2026
Abstract
(1) Background: Information diffusion models developed for graph-based platforms such as Reddit and broadcast architectures such as Telegram identify temporal features—particularly the timing of peak spread—as dominant predictors of coverage. Whether these predictors generalise to platforms where content is distributed through algorithmic recommendation [...] Read more.
(1) Background: Information diffusion models developed for graph-based platforms such as Reddit and broadcast architectures such as Telegram identify temporal features—particularly the timing of peak spread—as dominant predictors of coverage. Whether these predictors generalise to platforms where content is distributed through algorithmic recommendation rather than social-graph contagion remains an open question. (2) Methods: We analyse the YouNiverse dataset, comprising 133,364 English-language YouTube channels observed weekly from January 2015 to September 2019 (18.9 million observations). We derive channel-level diffusion features—including time-to-peak, post-peak decay rate, diffusion volatility, and upload frequency—and train three machine learning models (Linear Regression, Random Forest, and LightGBM) on two tasks: predicting peak weekly view growth (regression) and identifying viral channels (classification). A single-feature naive baseline (subscriber count alone) establishes the marginal contribution of the broader feature set beyond subscriber count alone, and a temporal split experiment (training on channels peaking before 2018, testing on 2018–2019) assesses cross-temporal stability. Because subscriber count and subscriber rank are measured at the October 2019 crawl, this is a retrospective characterisation rather than a strict real-time forecasting design. (3) Results: LightGBM achieves R2=0.776 (5-fold CV: 0.778±0.003) compared with R2=0.548 for the naive baseline, a net gain of +0.228R2. Because subscriber rank and subscriber count are near-perfectly collinear, we interpret them jointly as a channel-size dimension (42.2% of total mean absolute SHAP attribution), rather than as independent effects. Time-to-peak ranks fourteenth (1.1%), in contrast to its dominant role on Reddit (r=0.995, rank #1). For virality classification, LightGBM achieves ROC-AUC =0.967. Under the temporal split, Random Forest (R2=0.703) outperforms LightGBM (R2=0.683), showing greater cross-temporal stability within this retrospective split. (4) Conclusions: Within the 2015–2019 data, the results are consistent with algorithmic recommendation weakening the relationship between temporal diffusion dynamics and coverage magnitude at the channel level. Time-to-peak is weakly informative in this setting, while generalisation to the current recommendation system requires validation on newer data. Full article
(This article belongs to the Section Learning)
17 pages, 22828 KB  
Article
Spatiotemporal Dynamics and Potential Drivers of Cropland Fragmentation in the Yangtze River Delta, China, from 2000 to 2020
by Dongjie Li, Weiyang Chen and Bin Fang
Land 2026, 15(9), 1651; https://doi.org/10.3390/land15091651 (registering DOI) - 6 Sep 2026
Abstract
Recent research has advanced fine-scale mapping and driver analysis of cropland fragmentation, but composite indices may mask structurally different fragmentation configurations, and evidence on how terrain and urban-system location jointly relate to fragmentation remains limited in rapidly urbanizing delta regions. This study quantified [...] Read more.
Recent research has advanced fine-scale mapping and driver analysis of cropland fragmentation, but composite indices may mask structurally different fragmentation configurations, and evidence on how terrain and urban-system location jointly relate to fragmentation remains limited in rapidly urbanizing delta regions. This study quantified cropland fragmentation in the Yangtze River Delta (YRD), China, between 2000 and 2020. A Cropland Fragmentation Index (CFI) integrating edge density (ED), patch density (PD), and mean patch area (MPA) was calculated at a 1 km grid scale, and K-means clustering was used to identify fragmentation configurations. Pearson correlation and random-forest regression were used to examine spatial associations with selected 2020 natural and socioeconomic variables. Cropland area declined by 7.97%, while the regional-mean CFI increased from 0.29 to 0.32. Four configurations were identified, with the largest type (38.45% of grids) characterized by small patches and complex boundaries. Elevation and slope showed the strongest bivariate correlations with CFI, whereas distance to urban areas had the highest random-forest importance. These results reveal distinct fragmentation pathways and support differentiated cropland management in rapidly urbanizing regions. Full article
(This article belongs to the Topic Food Security and Healthy Nutrition)
Show Figures

Figure 1

22 pages, 11230 KB  
Article
Physics-Informed Decoupled Machine Learning for Context-Aware EV Range Optimization and Multi-Objective Driver Advisory
by Maksymilian Mądziel and Tiziana Campisi
Energies 2026, 19(17), 4209; https://doi.org/10.3390/en19174209 (registering DOI) - 6 Sep 2026
Abstract
Auxiliary heating, ventilation, and air conditioning (HVAC) systems can reduce electric vehicle (EV) driving range by over 20%, yet prevailing machine learning estimators often suffer from temporal data leakage, uninterpretable black-box structures, and lack real-time driver feedback. To address these challenges, this study [...] Read more.
Auxiliary heating, ventilation, and air conditioning (HVAC) systems can reduce electric vehicle (EV) driving range by over 20%, yet prevailing machine learning estimators often suffer from temporal data leakage, uninterpretable black-box structures, and lack real-time driver feedback. To address these challenges, this study presents a physics-informed decoupled machine learning framework integrated with a multi-objective Pareto Human–Machine Interface (HMI) advisory system. Powertrain traction power is estimated using a HistGradientBoosting regressor incorporating a mechanistic Vehicle Specific Power (VSP) feature, while cabin thermal dynamics are modeled via a regularized Random Forest regressor enriched with a Newtonian thermal decay function. Evaluated across an empirical 55-trip dataset using a 5-Fold GroupKFold cross-validation protocol, the traction and thermal models achieved out-of-sample accuracy of R2 = 0.9869 (MAE = 0.71 kW) and R2 = 0.8656 (MAE = 0.25 kW), respectively. Feature attributions were verified using SHAP analysis. An onboard Pareto optimization loop dynamically balances range extension against passenger thermal discomfort to deliver actionable driver recommendations. Multi-trip evaluation indicates that a representative 30% auxiliary load suppression yields average net energy savings of 5.21% entirely through software-driven guidance. Full article
Show Figures

Figure 1

35 pages, 1536 KB  
Article
The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift
by Istiaque Bhuiyan, Rahma Mirza, Ariful Hoque and Tanvir Bhuiyan
FinTech 2026, 5(3), 77; https://doi.org/10.3390/fintech5030077 (registering DOI) - 6 Sep 2026
Abstract
This study develops an empirical audit framework for assessing the explainability–reliability gap in fraud detection: whether stable model explanations remain consistent with performance-based feature reliance under distribution shift. Using the Bank Account Fraud dataset suite, including a Base dataset and five biased variants, [...] Read more.
This study develops an empirical audit framework for assessing the explainability–reliability gap in fraud detection: whether stable model explanations remain consistent with performance-based feature reliance under distribution shift. Using the Bank Account Fraud dataset suite, including a Base dataset and five biased variants, the study examines group-size disparity, fraud-prevalence disparity, separability bias, and temporal shift. Logistic Regression, Linear SVC, and Random Forest are benchmarked using standard classification metrics, followed by cross-variant evaluation with Logistic Regression as the interpretable baseline. SHAP is used to assess explanation stability, while permutation importance measures performance-based feature reliance. The results show that accuracy and ROC-AUC can overstate practical effectiveness under severe class imbalance; notably, Random Forest retained useful discrimination while producing near-zero recall at the evaluated threshold. SHAP feature rankings remained relatively stable across variants, particularly for address-history, identity-similarity, credit-risk, device, and behavioral variables. However, permutation importance revealed weaker and more variable reliance on several SHAP-ranked features. The limited agreement between the two measures indicates a partial explainability–reliability gap. The findings show that explanation stability alone is insufficient for evaluating trustworthy fraud detection models and should be complemented by performance-based validation under biased and shifted deployment conditions. Full article
(This article belongs to the Special Issue FinTech and Financial Stability: Opportunities and Risks)
Show Figures

Figure 1

39 pages, 35712 KB  
Article
Quantifying Landslide Damage Characteristics and Influencing Factors Across Mountain Forest Watershed Types Using Random Forest and SHAP
by Jaejeong Kim and Dongyeob Kim
Forests 2026, 17(9), 1063; https://doi.org/10.3390/f17091063 (registering DOI) - 6 Sep 2026
Abstract
Landslide damage in mountain forest watersheds may vary not only with local slope conditions but also with hydrological and geomorphological connectivity within watersheds. This study quantitatively analyzed differences in landslide damage characteristics and influencing factors among forest watershed types to support watershed-based landslide [...] Read more.
Landslide damage in mountain forest watersheds may vary not only with local slope conditions but also with hydrological and geomorphological connectivity within watersheds. This study quantitatively analyzed differences in landslide damage characteristics and influencing factors among forest watershed types to support watershed-based landslide damage mitigation. We extracted 1809 landslide damage polygons that overlapped mountain forest watersheds in the Chungcheong region of the Republic of Korea during 2022–2024 and classified them into catchment, slope, and independent zones. Damage area, perimeter, width, length, and elongation ratio were compared among watershed types. Area, perimeter, width, and length were largest in catchment zones and smallest in independent zones. Random forest models were developed and evaluated using repeated stratified 5-fold cross-validation with 13 topographic, soil, forest, geological, and watershed morphometric factors, and SHAP analysis was applied to interpret factor contributions. Watershed area was identified as a common important factor across all watershed types. However, secondary key factors differed by type: relief, DBH class, and slope length were relatively important in catchment zones; slope gradient, DBH class, and relief in slope zones; and relief and slope gradient in independent zones. These findings indicate that landslide damage characteristics and landslide-influencing factors vary by forest watershed type and can support differentiated forest disaster management and landslide mitigation strategies. Full article
(This article belongs to the Section Natural Hazards and Risk Management)
Show Figures

Figure 1

29 pages, 12997 KB  
Article
Cross-Estuary Generalization of Color Front Identification Using DenseNet-121
by Yumeng Tian, Luanbin Yin, Wenzhou Wu, Peng Zhang and Huiping Jiang
J. Mar. Sci. Eng. 2026, 14(17), 1657; https://doi.org/10.3390/jmse14171657 (registering DOI) - 6 Sep 2026
Abstract
Remote identification of color fronts, defined as transition zones with sharp gradients in water optical properties, has suffered from non-transferable thresholds, severe areal over-detection, and weak cross-estuary generalization. To address these issues, we propose a framework that integrates multi-scale spectral–spatial features with DenseNet-121. [...] Read more.
Remote identification of color fronts, defined as transition zones with sharp gradients in water optical properties, has suffered from non-transferable thresholds, severe areal over-detection, and weak cross-estuary generalization. To address these issues, we propose a framework that integrates multi-scale spectral–spatial features with DenseNet-121. We constructed a 165-D vector, seven window scales (three × three to 15 × 15) × two statistical descriptors (means and standard deviations) × 11 bands + 11 bands, then rearranged it into a 3D tensor and resized it to a 2D image for DenseNet-121 transfer learning with red-band post-processing. On in-distribution tests, the model achieves 0.953 accuracy, 0.953 F1, outperforming random forest. Cross-estuary generalization yields a mean F1 (0.744). Performance varies with optical compatibility: the Mississippi (runoff-dominated) gives the best F1 (0.874), while the Pearl (multi-channel, runoff-tide co-controlled) drops to 0.607 due to heterogeneity and reversed reflectance patterns. The red-band constraint can help reduce areal false alarms and improve spatial coherence of frontal regions, but its effectiveness depends on optical separability. The output width reflects superposition of transition zone and window scale. We demonstrate the potential and boundary conditions of this approach for cross-estuary color front identification, offering insights for physically consistent and generalizable ocean color monitoring. Full article
(This article belongs to the Section Physical Oceanography)
Show Figures

Figure 1

25 pages, 102309 KB  
Article
TandemNet: A Multi-Scale Multiple-Instance Learning Framework for Early-Season Rice Yield Prediction
by Meiqi Zeng, Wenxi Wu, Ran Yang, Wanxin Zhang, Siya Du, Xingzhi Huang and Luo Liu
Remote Sens. 2026, 18(17), 3040; https://doi.org/10.3390/rs18173040 (registering DOI) - 5 Sep 2026
Abstract
Early-season crop yield prediction is critical for food security assessment and timely agricultural decision-making, yet large-scale remote-sensing applications are often constrained by a mismatch between pixel-level observations and county-level yield labels. This mismatch limits the use of fine-grained spatial heterogeneity, especially under partial-season [...] Read more.
Early-season crop yield prediction is critical for food security assessment and timely agricultural decision-making, yet large-scale remote-sensing applications are often constrained by a mismatch between pixel-level observations and county-level yield labels. This mismatch limits the use of fine-grained spatial heterogeneity, especially under partial-season observations. Single-scale approaches are also limited in capturing complementary pixel-level and county-level information, restricting representation of yield formation processes. To address this, we propose TandemNet, a multi-scale multiple-instance learning framework for early-season prediction of japonica rice yield in Northeast China. TandemNet treats each county–year as a bag of rice pixels and adopts a dual-branch architecture to jointly learn pixel-level growth trajectories and county-level statistical responses. A phenology-conditioned cross-attention module fuses the two scales under varying growing-season windows. Using Sentinel-1, Sentinel-2, and MODIS data from 2018 to 2023, leave-one-year-out validation shows that TandemNet outperforms Random Forest, XGBoost, LSTM, and Transformer baselines across most phenological stages. It achieves reliable prediction at the tillering stage, approximately 2–3 months before harvest, with an R2 of 0.69 and an RMSE of 655.98 kg/ha. Ablation and attention analyses further indicate a stage-dependent shift from local heterogeneity to county-level consistency. These results demonstrate that modeling cross-scale interactions improves the timeliness, accuracy, and interpretability of rice yield prediction. Full article
Show Figures

Figure 1

16 pages, 1495 KB  
Article
Machine Learning-Based Prediction of N2O Emissions from Tea Plantations and Identification of Driving Factors for Sustainable Nitrogen Management
by Xiaoting Jie, Xin Liu, Jianfei Sun, Yanqiu Huang, Jing Xu and Yuan Zeng
Sustainability 2026, 18(17), 9123; https://doi.org/10.3390/su18179123 (registering DOI) - 5 Sep 2026
Abstract
Tea plantations are high-input agricultural systems and have been recognized as hotspots of soil nitrous oxide (N2O) emissions; however, the key controlling factors of these emissions and their quantitative prediction remain insufficiently understood. We compiled 115 field-observation records from 26 published [...] Read more.
Tea plantations are high-input agricultural systems and have been recognized as hotspots of soil nitrous oxide (N2O) emissions; however, the key controlling factors of these emissions and their quantitative prediction remain insufficiently understood. We compiled 115 field-observation records from 26 published studies into a multi-factor database covering climate, soil properties, and fertilization management, and compared five machine learning models—multiple linear regression (MLR), ridge regression, support vector regression (SVR), random forest (RF), and gradient-boosting regression trees (GBRTs)—using 5-fold cross-validation, combined with Spearman correlation and feature-importance analyses. Annual N2O emissions varied widely (0.40–73.20 kg·hm−2·a−1; mean 9.85 kg·hm−2·a−1), and the mean direct emission factor (EFd, 2.04%) far exceeded the IPCC default value. Emissions were significantly positively correlated with total nitrogen (TN) input but negatively correlated with mean annual temperature (MAT) and mean annual precipitation (MAP). GBRT performed best, effectively capturing nonlinear multifactor interactions; TN input and soil pH were the dominant predictors, followed by rainfall. However, feature importance rankings were method-dependent: the RF/SHAP analysis ranked MAT first rather than fifth, reflecting the different algorithmic mechanisms of the two approaches. The GBRT-based model provides a useful tool for estimating tea-plantation N2O emissions (LOOCV R2 = 0.668) and quantitative support for sustainable nitrogen management and targeted greenhouse gas mitigation strategies in tea production systems. Full article
30 pages, 14091 KB  
Article
Machine Learning-Based GNSS Positioning Error Compensation for Static Receivers
by Viorel Carbune, Maria Gutu, Irina Cojuhari, Lilia Rotaru and Vladimir Melnic
Geosciences 2026, 16(9), 356; https://doi.org/10.3390/geosciences16090356 (registering DOI) - 5 Sep 2026
Abstract
Global Navigation Satellite Systems (GNSS) positioning accuracy is affected by multiple error sources, including atmospheric delays, multipath propagation, and receiver noise, which can significantly reduce positioning reliability in low-cost receivers. This study investigates the use of a feedforward neural network to compensate for [...] Read more.
Global Navigation Satellite Systems (GNSS) positioning accuracy is affected by multiple error sources, including atmospheric delays, multipath propagation, and receiver noise, which can significantly reduce positioning reliability in low-cost receivers. This study investigates the use of a feedforward neural network to compensate for positioning errors in a static GNSS receiver scenario. A synthetic dataset was generated in MATLAB/Simulink by simulating positioning perturbations around a known reference location. Consecutive coordinate differences were used as input features, and a compact feedforward neural network with 45 hidden neurons was trained using the Levenberg–Marquardt algorithm to estimate positioning error components. The proposed approach was evaluated through residual error distribution, regression, temporal dispersion, and spatial scatter analyses. The results indicate that, for the primary 10 m error scenario, neural network-based compensation reduced temporal dispersion by approximately 46% and produced a more compact spatial distribution of corrected positions around the reference location. The residual errors remained concentrated near zero, indicating improved positioning consistency under the investigated simulation conditions. Sensitivity analysis across nominal error radii of R95 = 1, 5, 10, 15, and 20 m showed consistent reductions in both RMSE and standard deviation for radii of 10 m and above, whereas no consistent improvement was observed at lower error levels. In a preliminary comparison with random forests, XGBoost, Long Short-Term Memory (LSTM), and Gated Recurrent Unit models using the same training, validation, and test samples, the Feedforward Neural Network (FNN) achieved competitive test MSE while requiring substantially less training time and runtime memory than the LSTM. These findings support the proof-of-concept feasibility of lightweight FNN-based correction for simulated static GNSS positioning. Future work will focus on validation using real GNSS measurements and extension to dynamic positioning applications. Full article
(This article belongs to the Special Issue Earth Observation by GNSS and GIS Techniques, 2nd Edition)
Show Figures

Figure 1

32 pages, 17968 KB  
Article
Evaluation of Morphometric Conditioning Factors and Antecedent Rainfall in the Occurrence of Torrential Flows in Colombian Andean Watersheds
by Laura Ortiz-Giraldo, Derly Gómez, Edwin F. García, Blanca A. Botero, Johnny Vega, Hernan Martinez-Carvajal and Edier Aristizábal
Water 2026, 18(17), 2201; https://doi.org/10.3390/w18172201 - 4 Sep 2026
Abstract
Torrential flows, a broad category of rapid hydrogeomorphic processes that in the Colombian Andes includes debris flows, mudflows, and hyperconcentrated flows, pose a major hazard in tropical mountain regions. This study used two complementary binary classification models to examine geomorphometric conditioning and antecedent [...] Read more.
Torrential flows, a broad category of rapid hydrogeomorphic processes that in the Colombian Andes includes debris flows, mudflows, and hyperconcentrated flows, pose a major hazard in tropical mountain regions. This study used two complementary binary classification models to examine geomorphometric conditioning and antecedent rainfall triggering of torrential flow occurrence. A 12.5 m ALOS PALSAR DEM and 42 years of daily rainfall data (1981–2023) from IDEAM rain gauges and CHIRPS v2 were analyzed in a GIS-based regional framework. Antecedent rainfall variables were aggregated at watershed scale using zonal statistics. The conditioning dataset comprised 642 watersheds (321 with documented events and 321 controls). Gradient boosting ranked first in the preliminary grouped holdout comparison, whereas the uncalibrated random forest achieved the highest mean score under spatial leave-one-province-out validation and was selected as the final conditioning model (mean ROC-AUC = 0.747 ± 0.052). Basin scale and relief were the leading morphometric associations. In the rainfall trigger model, previous day IDEAM mean rainfall and previous day IDEAM maximum rainfall were the two leading permutation importance predictors, followed by monthly CHIRPS rainfall; the 90-day IDEAM maximum accumulation ranked fourth. This ordering indicates that immediate rainfall dominated the fitted model, while longer antecedent wetness retained a secondary contribution. The results support watershed prioritization and regional hazard assessment; because operational rainfall thresholds were not derived, they should not be treated as a ready-to-use early-warning model. Full article
Show Figures

Figure 1

19 pages, 2351 KB  
Article
Software-Defined UHF RFID Asset Tracking in Metallic Aircraft Cabins via IMU-Assisted Adaptive Kalman Filtering and Distilled Edge Intelligence
by Melis Karadag and Ozgun Pinarer
Sensors 2026, 26(17), 5639; https://doi.org/10.3390/s26175639 - 4 Sep 2026
Abstract
Passive Ultra-High Frequency (UHF) Radio Frequency Identification (RFID) systems deployed in metallic commercial aircraft cabins suffer from severe multipath fading, non-stationary channel dynamics, and operator gait-induced signal jitter. Addressing these challenges without physical airframe modifications or regulatory recertification remains a critical operational bottleneck. [...] Read more.
Passive Ultra-High Frequency (UHF) Radio Frequency Identification (RFID) systems deployed in metallic commercial aircraft cabins suffer from severe multipath fading, non-stationary channel dynamics, and operator gait-induced signal jitter. Addressing these challenges without physical airframe modifications or regulatory recertification remains a critical operational bottleneck. This paper presents an edge-native, software-defined framework that integrates micro-electromechanical system (MEMS) inertial measurements with an IMU-assisted Adaptive Kalman Filter (AKF) and a distilled surrogate decision tree. The proposed algorithm extracts localized motion energy (EIMU) to dynamically scale the measurement noise covariance (Rk) prior to physical-layer signal corruption, thereby eliminating phase lag and power hunting. For deterministic edge execution on COTS handheld devices, surrogate model distillation compresses a parent Random Forest ensemble into an 8.2KB 13-leaf decision tree (depth 5) yielding 0.12ms inference latency. Empirical validation across 17 operational sessions in Airbus A320, Boeing 737, and Airbus A321 cabins (10,720 valid reads) demonstrates a 99.45% mean RSSI jitter reduction (95%CI:[99.21%,99.63%]) and a 7.30× suppression of transmit power oscillations. Statistically, asset detection completeness is fully preserved (0.791 vs. 0.795 baseline, z=0.281,p=0.779). Operating entirely within standard handheld software runtimes, this approach bypasses Supplemental Type Certificate (STC) requirements while ensuring robust aerospace asset visibility. Full article
33 pages, 3036 KB  
Article
Benchmarking Statistical Methods for Environmental Chemical Mixtures: Prediction, Interaction Detection, and an Applied Analysis of Metals, Essential Elements and Diabetes
by Aderonke Gbemi Adetunji and Emmanuel Obeng-Gyasi
Stats 2026, 9(5), 96; https://doi.org/10.3390/stats9050096 - 4 Sep 2026
Abstract
Background. Human populations are exposed to complex chemical mixtures, making interaction detection a central challenge in environmental epidemiology. We benchmarked methods for prediction and recovery of interaction structure. Methods. Eight approaches—main-effects Lasso (glmnet_main), interaction Lasso (glmnet_int), hierNet, Random Forests, Bayesian Kernel Machine [...] Read more.
Background. Human populations are exposed to complex chemical mixtures, making interaction detection a central challenge in environmental epidemiology. We benchmarked methods for prediction and recovery of interaction structure. Methods. Eight approaches—main-effects Lasso (glmnet_main), interaction Lasso (glmnet_int), hierNet, Random Forests, Bayesian Kernel Machine Regression (BKMR), quantile g-computation (qgcomp), weighted quantile sum regression (gWQS) and SuperLearner—were evaluated across eight linear/nonlinear, additive/interaction, continuous/binary data-generating processes (500 replicates each). Every method completed in all 500 replicates of all eight scenarios. Prediction was assessed on held-out test data using observed-outcome and oracle-referenced metrics; interaction detection was assessed against three known pairwise interactions among 45 candidate pairs, using both hard selection and a threshold-free ranking criterion. BKMR was evaluated at 2000 versus 25,000 MCMC iterations with multi-chain convergence diagnostics. Sensitivity analyses varied sample size, exposure correlation, signal strength, and interaction form. BKMR was also applied illustratively to six metals and prevalent diabetes in NHANES. Results. In additive settings, observed-outcome prediction was similar across methods, but oracle-referenced continuous-outcome error differed by up to six-fold. With interactions, interaction-aware methods clearly outperformed additive-only approaches on the continuous oracle-referenced metrics: in LMI, the oracle MSE was 1.57 for hierNet and 1.87 for glmnet_int against 3.80 for glmnet_main and 4.94 for qgcomp. glmnet_int and hierNet showed comparable sensitivity; hierNet had a modestly lower mean per-replicate false discovery proportion in paired comparisons, while pooled false discovery favored hierNet in the continuous scenarios and glmnet_int in the binary ones; pooled false discovery rates were 0.79 to 0.82 in every interaction scenario, so roughly four in five selected pairs were false. In the scenarios without true interactions, the pooled false discovery rate was exactly 1. Under threshold-free ranking, BKMR was competitive with the penalized methods (pair-ranking AUC: 0.758 to 0.781 across the four interaction scenarios). BKMR’s apparent instability at 2000 iterations reflected inadequate sampling: 93% of monitored parameters had a Gelman–Rubin statistic above 1.1 and the minimum effective sample size was 7.5, whereas at 25,000 iterations the median statistic was 1.02 and the oracle MSE in LMI fell from 6.38 to 2.08. In NHANES, lead, manganese, and iron had the highest posterior inclusion probabilities, with predominantly nonlinear exposure–response functions. Conclusions. Method choice matters most when interactions are present. Interaction-aware methods are preferable when joint effects are relevant, selected interactions require replication given the high false discovery burden, and BKMR comparisons should report sampling budgets and convergence diagnostics rather than treating a short chain as characteristic of the method. Full article
44 pages, 4233 KB  
Article
From Storm Damage Detection to Windthrow Susceptibility Mapping: Evaluating Regional Transferability in Radiata Pine Plantations
by Michael S. Watt, Andrew Holdaway, Sadeepa Jayathunga, Pete Watt, Kate Halstead and Tommaso Locatelli
Remote Sens. 2026, 18(17), 3020; https://doi.org/10.3390/rs18173020 - 4 Sep 2026
Abstract
Windthrow is a major disturbance risk for radiata pine (Pinus radiata D. Don) plantations, but operational susceptibility models must transfer across regions and storm events. We developed a multi-regional framework combining airborne laser scanning (ALS), aerial imagery, mapped stand and site variables, [...] Read more.
Windthrow is a major disturbance risk for radiata pine (Pinus radiata D. Don) plantations, but operational susceptibility models must transfer across regions and storm events. We developed a multi-regional framework combining airborne laser scanning (ALS), aerial imagery, mapped stand and site variables, climate, soils, and event-period weather. These data were used to detect storm damage and model windthrow susceptibility following major storm events in Gisborne, Hawke’s Bay and Tasman, New Zealand. Windthrow was mapped from repeat ALS canopy-height differencing in Gisborne and Hawke’s Bay, and from post-storm aerial imagery in Tasman, producing 29,244 balanced windthrow and no-windthrow plot observations. Random-forest models were evaluated using stand-grouped, spatially blocked and leave-one-region-out validation. Stand structure provided the strongest predictive signal, with windthrow concentrated in older, taller and higher-volume stands. Adding long-term climate produced the largest improvement beyond the Base stand/site formulation, giving a pooled ROC–AUC of 0.901 ± 0.013. Spatially blocked ROC–AUC for the selected model ranged across the three regions from 0.780 to 0.856, while leave-one-region-out ROC–AUC ranged from 0.649 to 0.802, demonstrating useful but region-dependent transfer. Adding soil and event-period weather did not consistently improve transferability. Prevalence-calibrated conditional scenario estimates increased with stand development under the mapped regional prevalence and the conditions represented by the reference events. These estimates provide a scalable basis for comparative windthrow-risk screening but should not be interpreted as independently validated absolute or annual windthrow probabilities. Full article
18 pages, 4440 KB  
Article
Quantitative Prediction of Coal–Gangue Content Using Terahertz Time-Domain Spectroscopy and Physics-Informed Machine Learning
by Zeping Liu, Lipeng Hu, Jianfei Xu, Yadong Yang, Sitong Li, Zhou Xu, Longhai Liu, Jiabao Li, Houli Liu and Dongdong Ye
Materials 2026, 19(17), 3776; https://doi.org/10.3390/ma19173776 - 4 Sep 2026
Abstract
Quantitative determination of gangue content is important for efficient coal use and intelligent coal–gangue separation. We combine transmission terahertz time-domain spectroscopy (THz-TDS), multidomain feature fusion, and machine learning to predict gangue mass fraction in coal–gangue mixtures. Time- and frequency-domain signals, refractive index, absorption [...] Read more.
Quantitative determination of gangue content is important for efficient coal use and intelligent coal–gangue separation. We combine transmission terahertz time-domain spectroscopy (THz-TDS), multidomain feature fusion, and machine learning to predict gangue mass fraction in coal–gangue mixtures. Time- and frequency-domain signals, refractive index, absorption and extinction coefficients, and complex permittivity were extracted from samples with different gangue contents. Five-fold cross-validation was used to compare random forest, support vector regression, Gaussian process regression, an artificial neural network, and an Effective Medium Theory-constrained Physics-Informed Neural Network (EMT-PINN). EMT-PINN achieved the best performance, with a coefficient of determination (R2) of 0.81 ± 0.15, a mean absolute error (MAE) 3.17 ± 0.59%, and a root mean square error (RMSE) of 5.79 ± 0.21%, compared with R2 values of 0.72 ± 0.08, 0.61 ± 0.21, 0.74 ± 0.11, and 0.64 ± 0.18 for RF, SVR, GPR, and ANN, respectively. These results demonstrate the potential of physics-informed THz spectroscopy for rapid and physically interpretable quantitative characterization of coal–gangue mixtures. Full article
27 pages, 2007 KB  
Article
Identification of Landslide Risks in the Subtropical Hilly Regions of Southern China Using Integrated Multi-Source Synthetic Aperture Radar Interferometry and Machine Learning
by Guanzhi Luo, Qinghua Zhan, Feiting Yi and Rui Chen
Appl. Sci. 2026, 16(17), 8820; https://doi.org/10.3390/app16178820 - 4 Sep 2026
Abstract
The subtropical hilly regions of southern China are characterized by dense vegetation and highly concealed landslides, making it difficult for traditional, single-source remote sensing methods to meet disaster prevention needs. The core scientific contribution of this study is the development of a hierarchical, [...] Read more.
The subtropical hilly regions of southern China are characterized by dense vegetation and highly concealed landslides, making it difficult for traditional, single-source remote sensing methods to meet disaster prevention needs. The core scientific contribution of this study is the development of a hierarchical, progressive hazard identification framework that bridges the gap between InSAR deformation detection and landslide risk identification. This study focuses on Mayang County, Hunan Province, China, and combines time-series InSAR data from C-band Sentinel-1 and L-band ALOS-2 with a random forest (RF) algorithm to construct an early-stage identification model for landslide risks. All SAR data were processed under controlled baseline conditions (perpendicular baseline <150 m; polarization: VV for Sentinel-1, HH for ALOS-2). By screening highly reliable deformation points through dual-source cross-validation and integrating nine evaluation factors including slope, we established a two-layer coupled identification model combining InSAR deformation and susceptibility indices at the slope unit scale. The results showed that the dual-source InSAR approach achieved an identification accuracy of 71% (precision 68%, recall 65%, F1-score 0.66, Cohen’s κ 0.62), significantly outperforming single-source methods (62% for Sentinel-1 alone and 58% for ALOS-2 alone); the AUC was 0.815 under spatial block cross-validation, with an out-of-bag error of 16.8%. The dual-source InSAR approach identified a total of 59 potential hazard sites, 83.1% of which were located in medium- to high-risk zones. Following field surveys and LiDAR verification, 35 of these were confirmed as active landslide sites, demonstrating identification accuracy significantly superior to that of a single data source. The multi-source coupling framework proposed in this study effectively overcomes the decoherence issues associated with single-source SAR data in subtropical vegetated areas, providing reliable technical support for the early identification of landslides in humid hilly regions of southern China. Full article
(This article belongs to the Section Earth Sciences)
Back to TopTop