Next Article in Journal
A Comparative Study of Large Language Models for Industrial Cyber-Physical Security
Next Article in Special Issue
Water Level Estimation by Means of Microwave Reflection Measurements and Machine Learning Processing in a Multimode Cavity
Previous Article in Journal
Deep Learning Image Steganography Based on Dual-Path Fusion in Frequency and Spatial Domains
Previous Article in Special Issue
BASOSH—A Conceptual Framework and Literature Review on Bodycentric Antenna Systems for Occupational Safety and Health
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Physics-Supported Linear and Nonlinear Dimensionality Reduction for Supervised Adaptive Channel Selection in Hybrid RF-FSO-THz Communication Systems

by
Luis Miguel Pires
1,2,3,* and
Vitor Fialho
2,4,*
1
Technologies and Engineering School (EET), Instituto Politécnico da Lusofonia (IPLuso), 1700-098 Lisbon, Portugal
2
Department of Electronical Engineering, Telecommunications and Computers (DEETC), Instituto Superior de Engenharia de Lisboa (ISEL), 1959-007 Lisbon, Portugal
3
School of Communication, Arts and Information (ECATI), Lusofona University, 1749-024 Lisbon, Portugal
4
UNINOVA-CTS, NOVA University of Lisbon, Campus de Caparica, 2829-516 Monte de Caparica, Portugal
*
Authors to whom correspondence should be addressed.
Electronics 2026, 15(13), 2778; https://doi.org/10.3390/electronics15132778
Submission received: 29 May 2026 / Revised: 22 June 2026 / Accepted: 23 June 2026 / Published: 24 June 2026

Abstract

Hybrid RF-FSO-THz communication systems are promising candidates for future Internet of Things (IoT) and 6G networks because they combine the robustness of radio frequency links, the high-capacity potential of Free-Space Optical communications, and the ultra-wideband capabilities of terahertz transmission. Adaptive channel selection in such systems depends on multiple correlated environmental and physical-layer variables, including distance, rain intensity, humidity, visibility, turbulence strength, signal-to-noise ratio, channel capacity, and energy-efficiency metrics. This paper presents a physics-supported benchmark framework for supervised adaptive channel selection in hybrid RF-FSO-THz systems and systematically investigates the impact of linear and nonlinear dimensionality-reduction techniques on predictive performance, statistical robustness, computational complexity, and physical interpretability. A multi-scenario dataset comprising 5000 samples was generated using calibrated RF, FSO, and THz propagation models under clear, rain, fog, and worst-case environmental conditions. Principal Component Analysis (PCA) and Kernel PCA were evaluated together with Random Forest, Support Vector Machines (SVMs), XGBoost, Gradient Boosting (GB), Multi-Layer Perceptron (MLP), Logistic Regression, and Decision Trees. The results demonstrate that PCA preserves nearly all predictive capabilities while reducing the original 33-dimensional feature space by approximately 81.8%, maintaining accuracies close to 97–98% with the best-performing classifiers. Statistical significance analysis confirms that PCA introduces only modest degradations, whereas Kernel PCA consistently reduces the predictive performance while increasing memory requirements and inference latency. Additional environmental-only validation experiments indicate that adaptive channel selection remains highly learnable even when only pre-selection environmental descriptors are available, partially mitigating concerns regarding self-consistency bias. Overall, the results suggest that PCA provides an advantageous compromise among predictive accuracy, computational efficiency, statistical robustness, and physical interpretability for supervised adaptive channel selection in physics-supported hybrid wireless communication systems.

1. Introduction

Future Internet of Things (IoT) and sixth-generation (6G) communication systems are expected to support highly reliable, high-capacity, low-latency, and energy-efficient services under increasingly heterogeneous and dynamic propagation environments [1]. These requirements are difficult to satisfy with a single communication technology, as each transmission medium exhibits distinct advantages and limitations. Radio frequency (RF) communication [2] is generally more resilient under adverse environmental conditions and longer transmission distances, whereas Free-Space Optical (FSO) communication [3] offers very high-capacity optical links under clear atmospheric conditions. Terahertz (THz) communication [4,5] provides ultra-wideband short-range transmission with high directional gain, although it is highly sensitive to molecular absorption and humidity.
Hybrid RF-FSO-THz architectures exploit the complementary characteristics of these technologies by dynamically selecting the most appropriate transmission link according to environmental and channel conditions. In such systems, adaptive channel selection depends on a set of environmental variables, such as distance, rain rate, humidity, visibility, and turbulence, as well as physical-layer metrics such as the signal-to-noise ratio (SNR), capacity, outage tendency, and energy per bit. Since many of these variables originate from common propagation phenomena, they often exhibit strong physical correlations, resulting in substantial redundancy within the feature space.
Dimensionality-reduction techniques are therefore attractive for mitigating feature redundancy while reducing computational complexity. Principal Component Analysis (PCA) is a classical linear technique that projects the original variables into orthogonal components ordered by explained variance [6]. Kernel PCA extends this approach by applying PCA within an implicit nonlinear feature space through kernel functions [7,8]. However, although PCA and Kernel PCA have been widely applied in data analysis and pattern recognition, their specific impact on supervised adaptive channel selection in hybrid RF-FSO-THz systems remains insufficiently explored.
This paper extends dimensionality-reduction-based adaptive channel selection by introducing an explicit physical-layer modeling framework for RF, FSO, and THz propagation mechanisms. The generated communication metrics are derived from physically interpretable propagation equations, allowing the dimensionality-reduction analysis to be directly associated with propagation-aware communication variables. This paper extends dimensionality-reduction-based adaptive channel selection by introducing an explicit physical-layer modeling framework for RF, FSO, and THz propagation mechanisms. The generated communication metrics are derived from physically interpretable propagation equations, allowing dimensionality-reduction analyses to be directly associated with propagation-aware communication variables.
The primary contribution of this work is not the development of new machine-learning algorithms, but rather the construction of a physics-supported benchmark framework for adaptive channel selection in hybrid RF-FSO-THz communication systems. Within this framework, the impact of linear and nonlinear dimensionality-reduction techniques is systematically evaluated in terms of predictive performance, statistical robustness, computational efficiency, and physical interpretability.
The main contributions of this work are summarized as follows:
  • Development of a physics-supported RF-FSO-THz benchmark framework and propagation-aware dataset for supervised adaptive channel selection.
  • Comparative evaluation of linear (PCA) and nonlinear (Kernel PCA) dimensionality-reduction techniques across multiple supervised learning models.
  • Statistical, computational, and interpretability analyses of dimensionality reduction in adaptive channel-selection tasks.
  • Environmental-only validation and scenario-generalization experiments to assess robustness and mitigate potential self-consistency bias.
  • Comparison with physics-inspired analytical baselines derived from the underlying channel-selection criteria.

2. State of the Art and Related Work

RF communication remains the most established wireless technology and is widely adopted due to its robustness, coverage, and technological maturity [2]. Nevertheless, RF systems remain constrained by spectrum congestion and limited bandwidth. FSO communication has been proposed as a complementary solution because it enables high-capacity optical links without requiring a licensed spectrum [3]. However, FSO systems are highly sensitive to atmospheric attenuation, fog, scintillation, and pointing errors. THz communication is also considered a key enabling technology for future 6G systems because of its extremely large bandwidth. However, it suffers from molecular absorption and high propagation losses, particularly under humid conditions [4,5].
Hybrid wireless architectures have been proposed to overcome the limitations of individual communication technologies. In hybrid FSO/RF systems, RF links are commonly employed as backup channels when optical links are degraded by fog, turbulence, or pointing instability. More recent studies have explored RF/THz coexistence, FSO/THz integration, relay selection, and intelligent switching strategies. Nevertheless, many existing approaches remain limited to dual-technology configurations or rely on deterministic switching policies. Comparatively fewer studies jointly combine physics-based propagation modeling, dimensionality reduction, and supervised adaptive decision-making within a unified RF-FSO-THz framework.
Recent studies have increasingly explored intelligent adaptive switching mechanisms in hybrid optical and wireless communication systems. A broader survey of hybrid FSO/RF systems is presented in [9,10], covering propagation models, RF backup strategies, simulation frameworks, machine-learning approaches, practical deployment challenges, and future research directions. This survey highlights the growing relevance of hybrid architectures and emphasizes the increasing role of adaptive and learning-based mechanisms in heterogeneous wireless systems.
More recently, hybrid FSO/THz architectures with intelligent switching mechanisms have also been investigated. In [11], a seamless Gbps hybrid FSO/THz communication framework was proposed using intelligent switching to maintain high-capacity transmission under varying environmental conditions. Integrated THz/FSO systems are further reviewed in [12], where practical propagation constraints, atmospheric effects, deployment limitations, and 6G application scenarios are analyzed.
Machine-learning techniques are now widely applied to wireless link adaptation, intelligent switching, channel prediction, and resource allocation in hybrid communication systems [9,10,11,12]. Random Forest models are particularly effective for structured tabular datasets and nonlinear decision problems [13], whereas Support Vector Machines (SVMs) provide compact margin-based supervised classification baselines [14]. PCA is one of the most widely used techniques for reducing correlated numerical data while preserving most of the original variance [6]. Kernel PCA extends conventional PCA by replacing inner products with kernel evaluations, thereby enabling the representation of nonlinear structures in an implicit feature space [7,8]. Despite the growing adoption of machine learning in hybrid communication systems, comparatively little attention has been devoted to understanding how dimensionality-reduction techniques influence predictive performance, computational efficiency, statistical robustness, and physical interpretability. Table 1 summarizes the main research directions related to RF, FSO, THz, hybrid communication systems, dimensionality reduction, and supervised learning, while positioning the contribution of the proposed framework.
Existing studies primarily investigate hybrid architecture, intelligent switching, or machine learning independently. Comparative investigations combining dimensionality reduction, physical interpretability, computational cost analysis, and supervised adaptive channel selection in RF-FSO-THz systems remain relatively scarce. Recent measurement campaigns have further improved the understanding of mmWave and sub-THz propagation characteristics in realistic environments, including aircraft cabins and confined indoor scenarios [15]. In addition, programmable metasurfaces have recently emerged as promising enablers for beamforming, wavefront synchronization, and adaptive THz communication architectures [16].
Recent hardware-oriented advances further reinforce the relevance of adaptive hybrid communication architectures. Digital-encoding and programmable metasurfaces enable reconfigurable wavefront control, beam steering, and dynamic propagation-environment shaping, which may support future adaptive RF-FSO-THz links. In parallel, photonics-assisted THz communication technologies are emerging as promising solutions for high-frequency signal generation, broadband transmission, and integration with optical wireless systems. These developments suggest that future adaptive channel-selection frameworks may operate jointly with reconfigurable meta surface-assisted environments and photonics-enabled THz front ends [17,18].
Overall, the literature demonstrates growing interest in hybrid RF, FSO, and THz communication systems, intelligent switching mechanisms, and machine-learning-assisted adaptation. However, the combined evaluation of dimensionality reduction, computational efficiency, physical interpretability, and supervised adaptive channel selection within a unified physics-supported RF-FSO-THz framework remains insufficiently explored. This gap motivates the investigation presented in this work.

3. Materials and Methods

This section presents the proposed physics-supported RF-FSO-THz adaptive channel-selection framework, including environmental scenario generation, propagation modeling, link-level metric computation, dataset construction, dimensionality reduction, and supervised learning procedures.

3.1. Overall Framework

The proposed framework (Figure 1) follows a physics-to-learning workflow in which environmental conditions are first generated, followed by RF, FSO, and THz propagation modeling, computation of communication metrics, best-channel assignment, dimensionality reduction, and supervised classification.

3.2. Environmental Scenarios

Four scenarios are considered:
S = { clear ,   rain ,   fog ,   worst }
Each sample is described by
e i = [ d i , R i , H i , V i , C n , i 2 ]
where d i is distance, R i is rain rate, H i is relative humidity, V i is visibility, and C n , i 2 is the refractive-index structure parameter. Scenario definitions follow calibrated ranges for rain, humidity, visibility, and turbulence, resulting in a dataset containing 5000 total samples, with 1250 samples per scenario. The environmental conditions considered throughout the dataset generation process are summarized in Table 2.
RF communication remains comparatively robust under severe atmospheric degradation and over long transmission distances, although it is constrained by limited spectral efficiency. FSO links can provide very high-capacity transmission under favorable visibility conditions, but their performance degrades significantly in the presence of fog and atmospheric turbulence. THz communication offers ultra-wide bandwidth and high directional gain for short-range links; however, its performance is strongly influenced by humidity-induced molecular absorption and propagation losses.
Table 3 summarizes the main physical and simulation parameters adopted throughout the propagation modeling and dataset generation process.

3.3. RF Propagation Model

RF propagation model (Figure 2) combines large-scale free-space attenuation, rain-induced attenuation, and stochastic shadowing effects. This simplified model captures the dominant propagation impairments typically observed in sub-6 GHz and microwave terrestrial wireless channels. RF link is modeled using free-space path loss, rain attenuation, and log-normal shadowing. Total RF loss is
L R F , d B = L F S P L , d B R F + L r a i n , d B + L s h a d o w , d B
Free-space path loss (FSPL) is
L F S P L , d B R F = 20 l o g 10 4 π d f R F c 0
where d is distance, f R F is carrier frequency, and c 0 is speed of light [2]. Rain attenuation follows ITU-R power-law model [19]:
γ R = k R α ,
L r a i n , d B = γ R d k m .
Log-normal shadowing term is
L s h a d o w , d B N ( 0 , σ s h a d o w 2 ) .
Overall, RF formulation combines free-space attenuation, rain-induced loss, and stochastic shadowing effects.

3.4. FSO Propagation Model

FSO model (Figure 3) incorporates the main physical mechanisms affecting optical wireless propagation, namely, geometric spreading, atmospheric attenuation caused by visibility degradation, and turbulence-induced scintillation effects. FSO channel combines optical geometric loss, atmospheric attenuation, and scintillation effects:
L F S O , d B = L g e o m , d B F S O + L a t m , d B + L s c i n t , d B .
Optical geometric loss is
L g e o m , d B F S O = 20 l o g 10 4 π d λ G o p t ,
where λ is optical wavelength, and G o p t is effective optical aperture gain. Atmospheric attenuation follows Beer–Lambert law [3]:
L a t m , d B = 4.343 σ d k m .
Extinction coefficient is computed using a Kruse-type visibility model:
σ = 3.912 V λ 550   [ n m ] q
Turbulence is modeled using Rytov variance [20]:
σ R 2 = 1.23 C n 2 k o 7 / 6 d 11 / 6 ,
with
k o = 2 π λ .
Scintillation penalty is approximated by
L s c i n t , d B 2 σ R .
Overall, the FSO formulation captures visibility-dependent attenuation and turbulence-induced scintillation effects.

3.5. THz Propagation Model

THz propagation is primarily dominated by free-space attenuation, humidity-dependent molecular absorption, and highly directional antenna gains. THz channel (Figure 4) includes free-space path loss, gas absorption, and directional antenna gain:
L T H z , d B = L F S P L , d B T H z + L g a s , d B G T H z .
THz free-space path loss is
L F S P L , d B T H z = 20 l o g 10 4 π d f T H z c 0 .
To maintain a computationally efficient simulator suitable for large-scale synthetic dataset generation, the gas absorption coefficient was approximated using a first-order humidity-dependent model:
γ g a s = a 0 + a 1 H ,
where H denotes relative humidity.
This simplified formulation was adopted to emulate the dominant influence of atmospheric water vapor on THz propagation and should be interpreted as a computationally efficient engineering approximation inspired by ITU-R P.676-13 atmospheric absorption models [21] and previous THz propagation studies [22]. Although it does not explicitly model individual molecular absorption lines, the formulation captures the primary humidity-dependent attenuation trends required for large-scale benchmark generation while maintaining tractable simulation complexity. More detailed spectroscopic approaches based on HITRAN databases [23] or line-by-line molecular absorption models may be considered in future work to improve propagation realism and support experimental validation.
Gas absorption loss is
L g a s , d B = γ g a s d k m .
Consequently, THz communication is highly sensitive to both humidity and transmission distance.

3.6. Link Metrics and Best-Channel Labeling

After computing propagation losses for each communication technology, several physical-layer link metrics are derived, including received power, noise power, signal-to-noise ratio, channel capacity, and energy per transmitted bit. These metrics are subsequently used to define the adaptive channel-selection target employed during supervised learning. For each channel c { R F , F S O , T H z } , received power is
P r , c , d B m = P t , c , d B m L c , d B
Noise power is
N c , d B m = N 0 + 10 l o g 10 ( B c ) + N F c ,
where N 0 = 174   d Bm/Hz, B c is bandwidth, and N F c is noise figure.
SNR is
S N R c , d B = P r , c , d B m N c , d B m .
In linear scale,
S N R c = 10 S N R c , d B / 10 .
Shannon capacity is
C c = B c l o g 2 ( 1 + S N R c ) .
Energy per bit is
E b , c = P t , c C c .
Target is
y i { R F , F S O , T H z }
Two possible labeling strategies exist. The first is a pure maximum-SNR rule:
y i = a r g m a x c { R F , F S O , T H z } S N R c
In the present study, the feasibility thresholds were fixed as SNRmin = 0 dB and Cmin = 1 Mbps, representing conservative minimum communication requirements adopted throughout the dataset generation process.
However, a more physically meaningful formulation is based on a constrained energy-aware rule:
C v a l i d = { c : S N R c , d B S N R m i n C c C m i n }
Selected channel is
y i = a r g m i n c C v a l i d E b , c
If no channel satisfies the feasibility constraints, RF is selected as a robust fallback option:
y i = R F , C v a l i d =
In this formulation, SNR primarily determines communication feasibility, whereas energy per bit is used to perform the final channel-selection decision among feasible candidates.

3.7. Dataset Generation and Feature Matrix

Figure 5 summarizes the physics-supported dataset generation workflow adopted in this study. For each environmental scenario, the simulator generates samples of the physical variables, applies the RF, FSO, and THz channel models, computes corresponding link-level metrics, and assigns the best-channel label used in the supervised learning stage.
Each sample is represented as
x i = [ d i , R i , H i , V i , C n , i 2 , S N R R F , i , S N R F S O , i , S N R T H z , i , C R F , i , C F S O , i , C T H z , i , E b , R F , i , E b , F S O , i , E b , T H z , i ]
Full matrix is
X R N × D
All features are standardized:
x ~ i j = x i j μ j σ j .
Standardization is required because both PCA and SVM are sensitive to the feature scale [6,14].
The full-feature configuration includes both environmental descriptors and derived channel-quality metrics, such as path loss, SNR, capacity, and energy per bit. Since some of these derived metrics are also involved in the channel-labeling rule, this configuration represents a controlled physics-supported benchmark. To distinguish between variables available before channel selection and variables directly related to label generation, an additional environmental-only feature configuration was evaluated in Section 4.9. This environmental-only configuration was introduced to assess whether adaptive channel selection remains learnable using only variables available prior to channel selection and to partially mitigate concerns regarding self-consistency bias.

3.8. PCA

PCA is applied to the standardized feature matrix to obtain a compact lower-dimensional representation of the original feature space. PCA was selected because the generated feature space contains multiple physically related variables, including attenuation, received power, SNR, capacity, and energy-per-bit metrics, which are expected to exhibit substantial correlation and redundancy. PCA is applied to X ~ . Covariance matrix is
Σ = 1 N 1 X ~ T X ~
Eigenvalue problem is
Σ w j = λ j w j
The reduced representation with k components is given by
Z P C A = X ~ W k
where
W k = [ w 1 , w 2 , , w k ]
The Explained Variance Ratio (EVR) quantifies the proportion of the total dataset variance captured by each principal component and is used to determine the number of retained components given by
E V R ( k ) = j = 1 k λ j j = 1 D λ j
In the proposed RF-FSO-THz framework, PCA is expected to identify a compact set of latent components that capture the dominant propagation-dependent variability while reducing feature redundancy and computational complexity.

3.9. Kernel PCA

Kernel PCA extends conventional PCA by implicitly projecting the data into a nonlinear feature space using kernel evaluations. Kernel PCA maps data into an implicit feature space:
ϕ : R D F
Kernel matrix is
K i j = k ( x i , x j ) = ϕ ( x i ) , ϕ ( x j )
Using a Radial Basis Function (RBF) kernel [7]:
k ( x i , x j ) = e x p ( γ x i x j 2 )
RBF kernel was selected because it can effectively capture smooth nonlinear relationships between environmental variables and channel performance metrics while preserving local neighborhood structures. In hybrid RF-FSO-THz systems, propagation phenomena such as atmospheric absorption, turbulence effects, humidity variations, and visibility-dependent attenuation introduce nonlinear dependencies that may not be adequately represented through purely linear transformations. Among the commonly adopted kernel functions, the RBF kernel provides a flexible nonparametric representation that does not require prior assumptions regarding the form or degree of the underlying nonlinearity [7]. In the present study, Kernel PCA is included as a representative nonlinear dimensionality-reduction baseline intended for comparative analysis rather than as a fully optimized nonlinear embedding framework.
Centered kernel matrix is
K ~ = H K H
where
H = I 1 N 11 T
Kernel PCA solves the following eigenvalue problem:
K ~ α j = λ j α j
Reduced representation is
Z K P C A = [ α 1 , α 2 , , α k ]
Within the proposed RF-FSO-THz benchmark, Kernel PCA is evaluated to determine whether nonlinear feature transformations provide advantages over conventional linear PCA in representing propagation-driven channel-selection patterns.

3.10. Supervised Classification

All machine-learning models, dimensionality-reduction methods, cross-validation procedures, and evaluation metrics were implemented in Python 3.10 using scikit-learn [24]. Complete source code and reproducibility material are available in the project repository [25]. The generated datasets are evaluated using supervised classifiers to determine whether the reduced-dimensional representations preserve the discriminative structure required for adaptive channel selection. The objective is not to maximize predictive performance for a specific classifier, but rather to assess whether dimensionality-reduced feature representations preserve the discriminative structure required for adaptive channel selection.
Three feature configurations are evaluated:
X o r i g , Z P C A , Z K P C A
Classification problem is
y ^ i = f z i
Random Forest and SVMs [13,14] were selected as primary reference classifiers because they represent two complementary learning paradigms: ensemble tree-based learning and margin-based classification. Both methods have been widely adopted in communication-system optimization and provide strong baselines for evaluating dimensionality-reduction strategies. To further assess the robustness and generalizability of the obtained conclusions, additional experiments were conducted using Gradient Boosting [26], XGBoost [27], Multi-Layer Perceptrons [28], Decision Trees [29], and Logistic Regression [30], representing ensemble, neural, probabilistic, and tree-based learning paradigms. Accuracy is
A c c u r a c y = 1 N i = 1 N 𝟙 ( y ^ i = y i )
For class c :
P r e c i s i o n c = T P c T P c + F P c
R e c a l l c = T P c T P c + F N c
F1-Score is
F 1 c = 2 P r e c i s i o n c R e c a l l c P r e c i s i o n c + R e c a l l c
Macro-F1 is
F 1 m a c r o = 1 3 c { R F , F S O , T H z } F 1 c
Balanced accuracy is
B A = 1 3 c { R F , F S O , T H z } R e c a l l c
Additional experiments include repeated cross-validation, PCA sensitivity analysis, feature ablation, LOSO validation, cross-scenario transfer, and classifier consistency evaluation. Table 4 summarizes the experimental protocol adopted throughout the study, including dataset generation, feature configurations, classifiers, validation procedures, and evaluation metrics. All experiments were conducted using a fixed random seed to ensure reproducibility.

4. Results

This section presents the experimental evaluation of the proposed RF-FSO-THz dimensionality reduction framework, including correlation analysis, classifier comparison, PCA sensitivity analysis, feature ablation, and scenario-based generalization experiments.

4.1. Correlation Structure and Dataset Characteristics

The generated dataset combines environmental variables, propagation-dependent physical-layer metrics, and best-channel labels derived from the proposed adaptive RF-FSO-THz framework. Since the feature space includes multiple correlated physical variables, an initial statistical analysis was conducted to evaluate redundancy and feature interdependency before applying dimensionality-reduction techniques.
Figure 6 presents the Pearson correlation matrix of the generated variables. The observed correlation structure is physically expected because attenuation, received power, SNR, capacity, and energy-per-bit metrics are mathematically coupled through the underlying propagation equations. Environmental variables such as humidity, visibility, and rain intensity also exhibit strong interactions with the generated communication features. These correlations suggest the presence of substantial redundancy within the original feature space, thereby motivating the use of PCA-based dimensionality reduction.
Figure 7 illustrates the class distribution of the generated dataset. FSO represents the dominant class because it performs well under favorable atmospheric conditions, whereas RF and THz become preferable under specific propagation regimes. Although the dataset is moderately imbalanced, the class distribution remains sufficiently balanced to justify the use of macro-F1 and balanced accuracy as representative performance indicators.

4.2. Baseline Classifier Performance

Before evaluating dimensionality-reduction techniques, several baseline classifiers were trained using the original feature space. Table 5 presents the baseline supervised classification performance obtained using the original feature space prior to dimensionality reduction.
Random Forest achieved the highest overall performance, whereas the majority baseline produced substantially lower macro-F1 and balanced accuracy values. These results indicate that the generated dataset contains structured discriminative information consistent with the underlying propagation models. The strong performance achieved with Random Forest further suggests that the generated physical-layer dataset contains highly structured discriminative information.

4.3. Original Features Versus PCA and Kernel PCA

The main objective of this work is to evaluate the impact of linear and nonlinear dimensionality-reduction techniques on supervised adaptive channel selection. Table 6 summarizes the classification performance obtained using the original feature space, PCA-reduced representations, and Kernel PCA transformations under repeated stratified cross-validation.
PCA preserves most of the classification performance, particularly for Random Forest, despite the reduced dimensionality. Limited performance degradation suggests that the original feature space contains substantial redundancy that can be represented through a compact set of linear components. Kernel PCA, however, yields a consistently lower performance, especially for SVM. These results suggest that the dominant structure of the generated dataset is largely linear or near-linear.
Figure 8 presents a sensitivity analysis with respect to the number of retained PCA components.
The figure shows that classification performance rapidly increases up to approximately six principal components and then stabilizes. Most predictive capability is preserved after approximately six principal components, suggesting that the generated feature space contains substantial linear redundancy. From a physical perspective, the dominant principal components mainly capture joint variations associated with atmospheric attenuation, received power, and SNR-related communication quality indicators. Since several communication metrics are mathematically coupled through the propagation equations, PCA naturally concentrates most of the variance into a small number of dominant linear components.
The inferior performance of Kernel PCA suggests that the dominant class structure of the generated dataset is largely linear after feature standardization despite the nonlinear propagation mechanisms present in the underlying physical models.

4.4. PCA Projection and Decision Regions

To better understand the geometric structure of the dataset generated after dimensionality reduction, a two-dimensional PCA projection was analyzed. The preservation of class separability within the first two principal components provides additional evidence that the original feature space contains substantial redundancy.
Figure 9 illustrates decision regions obtained after projection into the two dominant principal components.
The projection reveals a relatively strong class separability even after reducing the original feature space to only two dimensions. RF samples mainly occupy negative PC1 regions, whereas FSO samples are more frequent in positive PC1 regions. THz samples mainly appear near intermediate transition regions associated with specific propagation conditions. The preservation of class structure within the reduced PCA domain explains why PCA maintains a classification performance close to that achieved using the original feature space.

4.5. Comparative Analysis of Original Features, PCA, and Kernel PCA

The primary objective of this analysis is not to identify the best classifier but rather to evaluate whether dimensionality reduction preserves the discriminative information required for adaptive channel selection.
Figure 10 summarizes the overall comparison between the original feature space, PCA-reduced representations, and Kernel PCA transformations.
The original feature space achieves the best overall performance. PCA remains close to the original representation, particularly for Random Forest, whereas Kernel PCA introduces a noticeable degradation, especially for the SVM. These results suggest that the generated communication dataset is largely governed by approximately linear propagation-dependent relationships. Overall, the results indicate that the adaptive RF-FSO-THz channel-selection problem is primarily governed by approximately linear relationships between environmental conditions and communication metrics.

4.6. Feature Ablation Analysis

To evaluate the relative contribution of different feature groups, a feature ablation study was conducted. Table 7 summarizes the impact of removing different feature groups from the original dataset during the feature ablation analysis.
Relatively limited degradation observed after removing individual feature groups suggests that adaptive decision information is distributed across multiple correlated predictors rather than concentrated in a single dominant variable. This observation further supports the effectiveness of PCA for reducing feature redundancy. The small effect sizes indicate that the performance differences between the original feature space and PCA representations are practically negligible despite statistical significance.

4.7. Leave One Scenario out Generalization (LOSO)

A Leave One Scenario Out (LOSO) validation strategy was adopted to evaluate generalization under unseen environmental conditions. Table 8 presents LOSO generalization results obtained under previously unseen environmental propagation conditions.
Figure 11 summarizes the LOSO generalization performance.
The model maintains a good generalization capability across unseen propagation scenarios. Nevertheless, rain and worst-case environments exhibit a slightly lower performance because of the increased variability and combined atmospheric impairments associated with these conditions.
Rain and worst-case scenarios exhibit larger variability because several propagation impairments occur simultaneously. In these environments, attenuation mechanisms associated with precipitation, reduced visibility, atmospheric turbulence, and molecular absorption jointly affect channel quality indicators, producing more complex and overlapping decision boundaries.
The computational results reinforce the practical attractiveness of PCA, which achieves a near-identical predictive performance while reducing computational requirements.

4.8. Cross-Scenario Transfer Performance

To further evaluate robustness, cross-scenario transfer experiments were performed by training classifiers on selected environmental conditions and testing them on previously unseen scenarios. Table 9 summarizes cross-scenario transfer results obtained when training and testing under different environmental scenario combinations.
The severe degradation observed in the final configuration can be physically explained by the absence of representative highly degraded optical propagation regimes during training. Fog and worst-case conditions simultaneously combine low visibility, high humidity, severe attenuation, and increased atmospheric uncertainty, leading to decision regions that cannot be adequately approximated from clear and moderate rain conditions alone. Although encouraging, these results should be interpreted within the context of the simulated propagation scenarios considered in this study.

4.9. Environmental-Only Validation

The previous experiments indicate that the generated dataset exhibits a highly structured decision space. However, since channel labels are derived from propagation-aware metrics, concerns regarding self-consistency bias may arise. To partially address this issue, an additional experiment was conducted using exclusively environmental descriptors, namely, transmission distance, rain intensity, relative humidity, visibility, and atmospheric turbulence strength.
All channel-dependent variables, including path loss, SNR, achievable capacity, and energy-per-bit metrics, were intentionally removed from the feature space. Several supervised classifiers were subsequently trained using the same repeated stratified cross-validation procedure adopted throughout the manuscript. Table 10 summarizes the results obtained.
The environmental-only experiment reveals that XGBoost, Gradient Boosting, and Random Forest maintained a balanced accuracy close to 0.97, suggesting that adaptive channel selection remains highly learnable, even from environmental conditions alone. Nevertheless, the generated dataset is still strongly constrained by the calibrated physics-supported simulator. Therefore, the environmental-only experiment should be interpreted as a controlled validation step designed to assess whether the adaptive channel-selection policy remains learnable from pre-selection environmental descriptors rather than as a representation of operational wireless deployments.
The strong performance obtained using only environmental variables demonstrates that adaptive channel-selection patterns remain largely predictable even after removing channel-dependent metrics directly associated with the labeling procedure.

4.10. Extended Classifier Benchmark

To assess whether the conclusions derived from Random Forest and SVMs remain valid across additional learning paradigms, further experiments were conducted considering Gradient Boosting, XGBoost, Multi-Layer Perceptrons, Logistic Regression, and Decision Trees.
Table 11 presents the predictive performance obtained using the original feature space, PCA representations, and Kernel PCA embeddings.
The obtained results reveal remarkable consistency among ensemble methods. Random Forest, Gradient Boosting, and XGBoost achieved an accuracy close to 98% when trained on the original feature space. PCA introduced only minor degradations, whereas the Kernel PCA systematically reduced predictive performance. These observations further support the hypothesis that the generated adaptive channel-selection problem is predominantly characterized by approximately linear structures. Interestingly, ensemble-based methods exhibited remarkable consistency, suggesting that the underlying adaptive channel-selection policy is highly structured and can be efficiently approximated through tree-based learning mechanisms.
Although the physics-inspired baseline achieves competitive performance, dimensionality-reduction-based supervised models provide a unified framework for feature compression, classifier comparison, robustness analysis, and future integration with measured datasets.

4.11. Kernel PCA Assessment and Statistical Significance Analysis

Although Kernel PCA theoretically enables nonlinear manifold extraction, the obtained results consistently indicate an inferior predictive performance compared with both the original representation and PCA. In this study, the RBF kernel was adopted owing to its widespread use in nonlinear dimensionality reduction. Kernel PCA was evaluated using six components and a fixed RBF parameter configuration to maintain comparability with the PCA experiments. A full hyperparameter search over γ values, component numbers, alternative kernels, and classifier-specific parameters was not performed. Consequently, the reported Kernel PCA results only reflect the adopted configuration and should not be generalized to all possible nonlinear embeddings. To determine whether the observed degradations are statistically meaningful, repeated stratified cross-validation experiments were performed using five folds and five repetitions. Paired t-tests, Wilcoxon signed-rank tests, and Cohen’s effect-size measures were subsequently computed. Table 12 presents the statistical significance analysis comparing the original feature space, PCA-reduced representations, and Kernel PCA embeddings using repeated stratified cross-validation (5 folds × 5 repetitions). Paired t-tests, Wilcoxon signed-rank tests, and Cohen’s d effect sizes are reported to quantify the magnitude and statistical relevance of the observed performance differences. Results indicate that PCA introduces relatively small but statistically significant reductions in predictive performance, whereas Kernel PCA leads to substantially larger degradations associated with very large effect sizes.
The results obtained demonstrate that, although PCA slightly reduces predictive performance, the degradation remains modest and statistically acceptable considering the significant reduction in dimensionality. In contrast, Kernel PCA consistently produces larger degradations, suggesting that the adaptive RF-FSO-THz channel-selection problem is predominantly characterized by approximately linear structures induced by the underlying physics-supported simulator.
The statistical analysis confirms that PCA introduces only modest but statistically significant reductions in predictive performance. In contrast, Kernel PCA produces considerably larger degradations associated with extremely large effect sizes. These findings suggest that the dominant discriminative structures induced by the underlying propagation simulator remain largely linear after standardization.

4.12. Computational Cost and Interpretability Analysis

One of the main motivations behind dimensionality reduction lies in reducing implementation complexity. Therefore, training time, inference latency, memory requirements, and dimensionality-reduction ratios were additionally evaluated. Table 13 summarizes the computational cost associated with the best-performing models.
PCA reduces dimensionality from 33 variables to six principal components, corresponding to an approximately 81.8% reduction. Despite this substantial compression, performance degradation remains limited to approximately 1–1.5%. This finding is particularly relevant for edge-AI applications, where memory footprint and inference latency are often more critical than marginal reductions in predictive accuracy.
Kernel PCA, however, requires approximately two orders of magnitude more memory while simultaneously increasing the inference latency.
Furthermore, PCA interpretability analysis revealed that the first principal component explains approximately 32.8% of the total variance and is mainly associated with RF and THz link-quality indicators. The second principal component predominantly captures atmospheric propagation conditions, including humidity, visibility, and turbulence-related variables. These observations indicate that PCA preserves physically meaningful structures rather than acting solely as a mathematical compression tool.
Table 14 provides PCA-explained variance and physical interpretation of the first principal components.

4.13. Physics-Inspired Analytical Baselines

Since the channel-selection labels are generated from a constrained energy-aware physical rule, analytical baselines were also evaluated. Three deterministic policies were considered: maximum-SNR selection, maximum-capacity selection, and minimum energy-per-bit selection among feasible links. The latter should be interpreted as a physics-inspired oracle baseline because it uses quantities closely related to the label-generation mechanism. Table 15 provides the physics-inspired analytical baseline performance.
The Min-Eb feasible rule achieves a performance comparable to the best supervised models, confirming that the generated labels are strongly associated with the underlying physical decision policy. In contrast, the Max-SNR and max-capacity rules perform substantially worse, showing that the adaptive channel-selection problem is not adequately represented by a single isolated metric. These results clarify that supervised models are mainly recovering the underlying energy-aware physical policy from the available features.

5. Discussion

The previous section presented the experimental evaluation of the proposed physics-supported dimensionality-reduction framework, including classifier comparisons, scenario-based validation, computational-cost assessment, and statistical significance analysis. The purpose of the present discussion is to interpret these findings from both communication-theoretic and machine-learning perspectives, emphasizing the implications of dimensionality reduction for adaptive RF-FSO-THz channel selection. Particular emphasis is placed on the influence of propagation-induced structures on supervised learning, the relative merits of PCA and Kernel PCA, the trade-offs between predictive performance and implementation complexity, and the potential integration of the proposed framework into future intelligent wireless architectures envisioned for beyond-5G and 6G systems.

5.1. Physics-Induced Structure of the Adaptive Channel Selection Problem

The relatively high classification performance observed throughout the experiments can be partially explained by the fact that the generated labels are derived from physically structured propagation models. Environmental variables directly influence attenuation, received power, SNR, achievable capacity, and energy efficiency. Consequently, the resulting decision boundaries exhibit substantial regularity and partial linear separability. This observation suggests that the supervised models are essentially learning to approximate a physics-supported adaptive channel-selection policy under controlled propagation conditions rather than extracting latent patterns from heterogeneous empirical measurements collected under operational conditions.
The environmental-only validation further supports this interpretation. Even after removing SNR, path loss, achievable capacity, and energy-per-bit metrics, the best-performing classifiers still achieved balanced accuracies close to 0.97. Nevertheless, it should be emphasized that the generated dataset remains strongly structured by the underlying calibrated propagation simulator. Therefore, the reported performance should be interpreted as a controlled benchmark rather than a direct representation of operational wireless deployments.
Furthermore, since channel-selection labels are generated from propagation-aware metrics, a degree of self-consistency is unavoidable. However, the environmental-only experiments indicate that the adaptive channel-selection policy can still be reasonably approximated by environmental descriptors alone, partially alleviating concerns regarding label leakage and self-consistency bias.

5.2. PCA Versus Kernel PCA

One of the most significant findings of this study concerns the effectiveness of PCA. Using only six principal components, Random Forest achieved an accuracy of approximately 0.97, compared with approximately 0.98 obtained using the original feature space. Similar trends were observed for XGBoost and Gradient Boosting.
These results confirm the existence of substantial redundancy in the generated dataset, which is physically expected since attenuation, received power, SNR, achievable throughput, and energy-per-bit metrics are mathematically interconnected. PCA effectively compresses this information while preserving most of the predictive capability.
In contrast, Kernel PCA consistently underperformed. Although RF, FSO, and THz propagation models inherently exhibit nonlinearities, the dominant class structure after standardization appears to predominantly be linear or near-linear. Kernel PCA may overemphasize local nonlinear variations and distort global geometric relationships among propagation states.
The PCA interpretability analysis further supports this interpretation. The first principal component explains approximately 32.8% of the total variance and is primarily dominated by THz and RF link-quality indicators, including path loss, SNR, and throughput. The second component is largely associated with environmental conditions, including humidity, visibility, and atmospheric turbulence strength. PCA preserves propagation-aware latent information rather than acting solely as a mathematical compression technique. Future work will investigate the systematic optimization of RBF kernel parameters, alternative kernel functions, and manifold-learning approaches such as Isometric Mapping (ISOMAP) and Uniform Manifold Approximation and Projection (UMAP).

5.3. Additional Classifiers and Statistical Analysis

The inclusion of additional machine-learning models demonstrates that Random Forest, Gradient Boosting, and XGBoost exhibit comparable predictive capabilities, all achieving accuracies close to 98%.
Statistical significance analysis based on repeated stratified cross-validation, paired t-tests, Wilcoxon signed-rank tests, and Cohen’s effect sizes confirms that PCA introduces only modest but statistically significant reductions in predictive performance. Conversely, Kernel PCA produces substantially larger degradations associated with extremely large effect sizes.
Taken together, these findings suggest that the adaptive channel-selection problem considered in this study is largely characterized by approximately linear structures resulting from the strong regularity imposed by the propagation-aware simulation framework.

5.4. Computational Efficiency

Dimensionality reduction is often motivated by implementation constraints in practical intelligent communication systems. The computational complexity analysis reveals that PCA provides an advantageous trade-off between predictive performance and implementation cost.
PCA reduces the dimensionality from 33 input variables to six principal components, corresponding to an approximately 81.8% reduction. At the same time, predictive degradation remains limited to approximately 1–1.5% for the best-performing classifiers.
Moreover, training times decrease noticeably, particularly for Gradient Boosting, where training time is reduced from approximately 12 s to 5 s. Inference times remain below 0.03 ms per sample for the best-performing models. Conversely, Kernel PCA requires approximately two orders of magnitude more memory (~189 MB versus ~2 MB) and exhibits considerably longer inference times while simultaneously degrading predictive performance.
These observations strongly suggest that PCA offers the most advantageous trade-off among compactness, robustness, interpretability, and computational efficiency.

5.5. Practical Deployment Considerations for IoT and 6G Systems

Although the proposed framework has been validated using a controlled physics-supported simulation environment, its architecture is compatible with intelligent communication systems envisioned for beyond-5G and 6G networks. The present propagation framework intentionally focuses on dominant first-order effects. Several practical impairments are not explicitly modeled, including pointing errors, beam divergence, aperture averaging, scintillation fading distributions, detailed molecular absorption spectra, hardware nonlinearities, modulation and coding schemes, relay architectures, and switching overheads. Consequently, the proposed framework should be interpreted as a controlled propagation-aware benchmark rather than a complete end-to-end communication system model.
In practical deployments, PCA can be trained offline using historical propagation data, whereas only the linear projection stage needs to be executed online. This significantly reduces computational requirements and makes the proposed approach suitable for deployment on edge-computing platforms, MEC infrastructures, intelligent gateways, or embedded controllers.
The measured inference times suggest that adaptive channel-selection decisions can be executed within latency budgets compatible with dynamic wireless environments. Furthermore, the framework may be integrated into AI-assisted radio resource management entities, near-real-time RAN Intelligent Controllers (near-RT RIC) within O-RAN architectures, digital-twin-assisted communication systems, and AI-native self-optimizing 6G networks.
Nevertheless, several challenges remain before practical deployment can realistically be envisaged. Real deployments involve temporal channel dynamics, hardware impairments, pointing instability, imperfect channel estimation, interference variability, user mobility, and measured propagation conditions. Future work will therefore focus on validation using experimental measurements, field trials, and publicly available hybrid communication datasets.
Future investigations should further assess robustness under perturbed propagation assumptions, unseen distance ranges, alternative humidity and visibility distributions, different carrier frequencies, modified bandwidths, and different SNR/capacity feasibility thresholds. Such experiments would help determine whether PCA-based compression remains stable beyond the specific synthetic dataset and parameter settings considered in this study. Although the present study focuses on supervised learning, the proposed benchmark may also serve as a suitable environment for future investigations involving reinforcement learning, contextual bandits, communication digital twins, or online adaptation mechanisms capable of continuously updating channel-selection policies under evolving propagation conditions. Overall, the obtained results indicate that PCA provides the most attractive compromise among predictive capability, physical interpretability, computational efficiency, and implementation feasibility, whereas Kernel PCA does not justify its additional complexity for the considered RF-FSO-THz adaptive channel-selection problem.
Beyond the specific RF-FSO-THz application considered in this work, the results suggest that physically structured communication datasets may often contain substantial redundancy arising from common propagation mechanisms. Consequently, interpretable linear dimensionality-reduction methods may provide a practical alternative to more complex nonlinear approaches when the dominant variability is governed by physically coupled communication metrics.

6. Conclusions

This paper presented a physics-supported benchmark framework for investigating linear and nonlinear dimensionality-reduction techniques in supervised adaptive channel selection for hybrid RF-FSO-THz communication systems. Rather than proposing novel machine-learning algorithms, the main contribution lies in systematically investigating how linear and nonlinear dimensionality-reduction techniques affect predictive performance, computational complexity, statistical robustness, and physical interpretability within a controlled propagation-aware environment.
The obtained results demonstrate that PCA preserves nearly all predictive capabilities while substantially reducing implementation complexity. Using only six principal components, Random Forest, Gradient Boosting, and XGBoost maintain accuracies close to 97–98% despite reducing the original 33-dimensional feature space by approximately 81.8%. Statistical significance analysis further confirms that PCA introduces only modest but statistically significant performance degradations, whereas Kernel PCA consistently produces larger reductions associated with very large effect sizes, increased inference latency, and significantly higher memory requirements.
Additional environmental-only validation experiments demonstrated that adaptive channel-selection patterns remain largely predictable even when channel-dependent metrics are excluded from the feature space.
Additional experiments reveal that adaptive channel selection remains highly learnable even when only environmental descriptors are available, partially mitigating concerns regarding self-consistency bias and label leakage. Furthermore, analytical physics-inspired baselines indicate that the minimum-energy feasible policy performs similarly to the best supervised models, suggesting that machine-learning approaches effectively recover the underlying energy-aware propagation policy encoded within the synthetic benchmark.
The interpretability analysis shows that the dominant principal components preserve physically meaningful latent structures associated with link quality, attenuation mechanisms, and atmospheric propagation conditions. Taken together, the obtained findings suggest that PCA provides the most advantageous compromise among predictive performance, physical interpretability, robustness, and computational efficiency for the considered RF-FSO-THz adaptive channel-selection problem. The reported results should be interpreted within the context of a controlled physics-supported benchmark generated from calibrated propagation models and simplified environmental assumptions.
Although the proposed framework relies on calibrated synthetic propagation models and intentionally focuses on dominant first-order effects, it provides a flexible benchmark for future investigations. Future work will focus on incorporating measured RF, FSO, and THz propagation data, more detailed atmospheric absorption models, temporal channel dynamics, hardware impairments, and deployment-oriented adaptive communication scenarios. Such extensions will enable further assessment of whether the dimensionality-reduction behavior observed in this study generalizes to real-world heterogeneous wireless environments.

Author Contributions

Conceptualization, L.M.P. and V.F.; methodology, L.M.P. and V.F.; software, L.M.P.; validation, L.M.P. and V.F.; formal analysis, L.M.P. and V.F.; investigation, L.M.P. and V.F.; resources, L.M.P.; data curation, L.M.P. and V.F.; writing—original draft preparation, L.M.P. and V.F.; writing—review and editing, L.M.P. and V.F.; visualization, L.M.P. and V.F.; supervision, L.M.P. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding authors.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
DTDecision Tree
EVRExplained Variance Ratio
FSOFree-Space Optical
FSPLFree-Space Path Loss
GBGradient Boosting
HITRANHigh-Resolution Transmission Molecular Absorption Database
IoTInternet of Things
ISOMAPIsometric Mapping
KPCAKernel Principal Component Analysis
LOSOLeave One Scenario Out
LRLogistic Regression
MECMulti-Access Edge Computing
MLMachine Learning
MLPMulti-Layer Perceptron
O-RANOpen Radio Access Network
PCAPrincipal Component Analysis
RANRadio Access Network
RBFRadial Basis Function
RFRadio Frequency
RICRAN Intelligent Controller
SNRSignal-to-Noise Ratio
SVMSupport Vector Machine
THzTerahertz
UMAPUniform Manifold Approximation and Projection
XGBExtreme Gradient Boosting

References

  1. Saad, W.; Bennis, M.; Chen, M. A vision of 6G wireless systems: Applications, trends, technologies, and open research problems. IEEE Netw. 2020, 34, 134–142. [Google Scholar] [CrossRef]
  2. Goldsmith, A. Wireless Communications; Cambridge University Press: Cambridge, UK, 2005. [Google Scholar]
  3. Khalighi, M.A.; Uysal, M. Survey on free space optical communication: A communication theory perspective. IEEE Commun. Surv. Tutor. 2014, 16, 2231–2258. [Google Scholar] [CrossRef]
  4. Jornet, J.M.; Akyildiz, I.F. Channel modeling and capacity analysis for electromagnetic wireless nanonetworks in the terahertz band. IEEE Trans. Wirel. Commun. 2011, 10, 3211–3221. [Google Scholar] [CrossRef]
  5. Chen, Z.; Ma, X.; Zhang, B.; Niu, J.; Kuang, L.; Li, C.; Wang, S.; Debbah, M. A survey on terahertz communications. China Commun. 2019, 16, 1–35. [Google Scholar] [CrossRef]
  6. Jolliffe, I.T.; Cadima, J. Principal component analysis: A review and recent developments. Philos. Trans. R. Soc. A Math. Phys. Eng. Sci. 2016, 374, 20150202. [Google Scholar] [CrossRef]
  7. Schölkopf, B.; Smola, A.; Müller, K.-R. Nonlinear component analysis as a kernel eigenvalue problem. Neural Comput. 1998, 10, 1299–1319. [Google Scholar] [CrossRef]
  8. Williams, C.K.I. On a connection between kernel PCA and metric multidimensional scaling. Mach. Learn. 2002, 46, 11–19. [Google Scholar] [CrossRef]
  9. Shao, J.; Liu, Y.; Du, X.; Xie, T. Adaptive Modulation Scheme for Soft-Switching Hybrid FSO/RF Links Based on Machine Learning. Photonics 2024, 11, 404. [Google Scholar] [CrossRef]
  10. Phuchortham, S.; Sabit, H. A Survey on Free-Space Optical Communication with RF Backup: Models, Simulations, Experience, Machine Learning, Challenges and Future Directions. Sensors 2025, 25, 3310. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  11. Gao, W.; Liu, K.; Han, C.; Chen, Z. Seamless Gbps Hybrid FSO/THz Communication with Intelligent Switching. J. Light. Technol. 2025, 43, 9079–9089. [Google Scholar] [CrossRef]
  12. Liu, J.; Yang, X.; Wei, Y.; Zhao, F. Integrated THz/FSO Communications: A Review of Practical Constraints, Applications and Challenges. Micromachines 2025, 16, 1297. [Google Scholar] [CrossRef] [PubMed]
  13. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
  14. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
  15. Liao, X.; Tang, H.; Zheng, X.; Hao, X.; Wang, Y. Channel Measurements, Characterization and Performance Analysis for an Aircraft Cabin From Mmwave to sub-THz Communication. IEEE Trans. Veh. Technol. 2026. [Google Scholar] [CrossRef]
  16. Zhang, C.; Lou, J.; Zhang, J.; Wang, Z.; Ji, C.; Yuan, H.; Huang, Y.; Li, Z.; Zhu, W.; Zhao, S.; et al. Space-Time Wavefront Synchronized Terahertz Metasurface. Adv. Mater. 2026, 38, e20890. [Google Scholar] [CrossRef] [PubMed]
  17. Zhan, H.; Gu, M.; Tian, Y.; Feng, H.; Zhu, M.; Zhou, H.; Jin, Y.; Tang, Y.; Li, C.; Fang, B.; et al. Review for wireless communication technology based on digital encoding metasurfaces. Opto-Electron. Adv. 2025, 8, 240315. [Google Scholar] [CrossRef]
  18. Tian, Y.; Dong, B.; Li, Y.; Xiong, B.; Zhang, J.; Sun, C.; Hao, Z.; Wang, J.; Wang, L.; Han, Y.; et al. Photonics-assisted THz wireless communication enabled by wide-bandwidth packaged back-illuminated modified uni-traveling-carrier photodiode. Opto-Electron. Sci. 2024, 3, 230051. [Google Scholar] [CrossRef]
  19. ITU-R. Recommendation ITU-R P.838-3: Specific Attenuation Model for Rain for Use in Prediction Methods; International Telecommunication Union: Geneva, Switzerland, 2005. [Google Scholar]
  20. Andrews, L.C.; Phillips, R.L. Laser Beam Propagation Through Random Media, 2nd ed.; SPIE Press: Bellingham, WA, USA, 2005. [Google Scholar]
  21. ITU-R. Recommendation P.676-13: Attenuation by Atmospheric Gases and Related Effects; International Telecommunication Union (ITU): Geneva, Switzerland, 2022. [Google Scholar]
  22. Han, C.; Wang, Y.; Li, Y.; Chen, Y.; Abbasi, N.A.; Kürner, T.; Molisch, A.F. Terahertz Wireless Channels: A Holistic Survey on Measurement, Modeling, and Analysis. IEEE Commun. Surv. Tutor. 2022, 24, 1670–1707. [Google Scholar] [CrossRef]
  23. Gordon, I.; Rothman, L.; Hargreaves, R.; Hashemi, R.; Karlovets, E.; Skinner, F.; Conway, E.; Hill, C.; Kochanov, R.; Tan, Y.; et al. The HITRAN2020 Molecular Spectroscopic Database. J. Quant. Spectrosc. Radiat. Transf. 2022, 277, 107949. [Google Scholar] [CrossRef]
  24. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
  25. Pires, L.M. Physics-Supported Dimensionality Reduction for Adaptive RF–FSO–THz Channel Selection. GitHub Repository. 2026. Available online: https://github.com/prof-luispires/Linear_Nonlinear_Dim_Reduction_Channel_Hybrid_RF_FSO_THz_-Comm.git (accessed on 6 May 2026).
  26. Friedman, J.H. Greedy Function Approximation: A Gradient Boosting Machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef]
  27. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
  28. Rumelhart, D.; Hinton, G.; Williams, R. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef]
  29. Breiman, L.; Friedman, J.H.; Olshen, R.A.; Stone, C.J. Classification and Regression Trees; Wadsworth International Group: Belmont, CA, USA, 1984. [Google Scholar]
  30. Hosmer, D.W.; Lemeshow, S.; Sturdivant, R.X. Applied Logistic Regression, 3rd ed.; Wiley: Hoboken, NJ, USA, 2013. [Google Scholar]
Figure 1. Proposed physics-supported dimensionality reduction and supervised channel-selection pipeline.
Figure 1. Proposed physics-supported dimensionality reduction and supervised channel-selection pipeline.
Electronics 15 02778 g001
Figure 2. RF channel model.
Figure 2. RF channel model.
Electronics 15 02778 g002
Figure 3. FSO channel model.
Figure 3. FSO channel model.
Electronics 15 02778 g003
Figure 4. THz channel model.
Figure 4. THz channel model.
Electronics 15 02778 g004
Figure 5. Physics-supported dataset generation pipeline for the hybrid RF-FSO-THz adaptive channel-selection dataset.
Figure 5. Physics-supported dataset generation pipeline for the hybrid RF-FSO-THz adaptive channel-selection dataset.
Electronics 15 02778 g005
Figure 6. Feature correlation heatmap of environmental and physical-layer variables used in the proposed adaptive RF-FSO-THz channel-selection framework.
Figure 6. Feature correlation heatmap of environmental and physical-layer variables used in the proposed adaptive RF-FSO-THz channel-selection framework.
Electronics 15 02778 g006
Figure 7. Distribution of best-channel labels in the generated RF-FSO-THz dataset.
Figure 7. Distribution of best-channel labels in the generated RF-FSO-THz dataset.
Electronics 15 02778 g007
Figure 8. Classification accuracy as a function of the number of retained PCA components for Random Forest, SVM, and Gradient Boosting classifiers.
Figure 8. Classification accuracy as a function of the number of retained PCA components for Random Forest, SVM, and Gradient Boosting classifiers.
Electronics 15 02778 g008
Figure 9. Decision regions obtained in a two-dimensional PCA projection of the RF-FSO-THz dataset.
Figure 9. Decision regions obtained in a two-dimensional PCA projection of the RF-FSO-THz dataset.
Electronics 15 02778 g009
Figure 10. Comparison between original features, PCA, and Kernel PCA for Random Forest and SVM classifiers using accuracy, macro-F1, and balanced accuracy metrics.
Figure 10. Comparison between original features, PCA, and Kernel PCA for Random Forest and SVM classifiers using accuracy, macro-F1, and balanced accuracy metrics.
Electronics 15 02778 g010
Figure 11. LOSO generalization performance across unseen environmental conditions.
Figure 11. LOSO generalization performance across unseen environmental conditions.
Electronics 15 02778 g011
Table 1. Related work and positioning of this paper.
Table 1. Related work and positioning of this paper.
AreaRepresentative ReferencesMain ContributionLimitation Addressed in This Paper
RF
communication
[2]Robust wireless propagation and link-budget modelingRF is integrated as one candidate in a hybrid adaptive system
FSO
communication
[3]Optical wireless links, visibility loss, turbulence limitationsFSO is modeled jointly with RF and THz
THz
communication
[4,5]Ultra-wideband links and molecular absorption effectsTHz is included as a third adaptive option
Hybrid FSO/RF systems[9,10]RF backup, soft-switching, adaptive modulation, and ML-assisted hybrid operationExtended here to a three-technology RF-FSO-THz adaptive framework
Hybrid FSO/THz systems[11,12]Intelligent switching and integrated THz/FSO architectures for high capacity 6G linksCombined here with supervised dimensionality reduction and multi-scenario classification
PCA[6]Linear feature compression and variance preservationPCA is evaluated for RF-FSO-THz channel selection
Kernel PCA[7,8]Nonlinear embedding through kernelsKernel PCA is compared directly against PCA
Supervised ML[13,14]Classification with Random Forest and SVMModels are evaluated under original, PCA, and Kernel PCA features
Proposed workThis paperPhysics-supported dimensionality reduction for adaptive RF-FSO-THz channel selectionIntegrates physical modeling, dimensionality reduction, supervised classification, and scenario
generalization
Table 2. Environmental scenario configuration used in the simulation.
Table 2. Environmental scenario configuration used in the simulation.
ScenarioRain Rate R Humidity H Visibility V Turbulence C n 2 Expected
Physical Effect
Clear 0 mm/h30–50%15–25 km 1 × 10 15 FSO/THz
favorable
Rain25–50 mm/h60–80%3–8 km 5 × 10 15 RF becomes more relevant
Fog5–25 mm/h80–95%0.3–1.0 km 1 × 10 14 FSO strongly
degraded
Worst60–80 mm/h90–100%0.1–0.4 km 5 × 10 14 RF dominance expected
Table 3. Simulation parameters and physical assumptions.
Table 3. Simulation parameters and physical assumptions.
ParameterRFFSOTHzDescription
Carrier
frequency
5 GHz193 THz (1550 nm)300 GHzOperating
frequency
Distance range10–4000 m10–4000 m10–4000 mCommunication distance
Transmit power20 dBm10 dBm0 dBmTransmission power
Antenna gain5 dBiOptical gain40 dBiDirectional gain
RF rain
attenuation
ITU-R P.838Rain model
Visibility range0.1–25 kmOptical attenuation
Humidity range50–100%Molecular
absorption
Turbulence   parameter ,   C n 2 10−16–10−13Atmospheric
turbulence
Noise bandwidth20 MHz1 GHz10 GHzEffective
bandwidth
Noise figure5 dB3 dB7 dBReceiver noise assumption
Table 4. Summary of the experimental configuration adopted throughout the study.
Table 4. Summary of the experimental configuration adopted throughout the study.
ComponentConfiguration
DatasetPhysics-supported simulated RF-FSO-THz
dataset
ScenariosClear, rain, fog, worst
Samples5000 samples
Random seedFixed random seed = 42
Sampling strategyUniform random sampling within each
scenario-specific environmental range
Train/test protocolStratified splits and repeated stratified
cross-validation
Target variableBest communication channel (RF, FSO, THz)
Environmental-only featuresDistance, rain rate, humidity, visibility,
turbulence strength
Original feature space33 environmental and communication features
Reduced PCA space6 principal components
Kernel PCARBF kernel
Primary classifiersRandom Forest, SVM
Additional classifiersXGBoost (XGB), GB, MLP, DT, Logistic
Regression (LR)
BaselinesMajority baseline, maximum-SNR rule,
maximum-capacity rule, minimum energy-per-bit feasible rule
Evaluation metricsAccuracy, Macro-F1, balanced accuracy
Statistical analysispaired t-test, Wilcoxon signed-rank test,
Cohen’s d
Cross-validationRepeated stratified K-Fold
(5 folds × 5 repetitions)
Generalization testsLeave one scenario out (LOSO),
cross-scenario transfer
Additional
experiments
Environmental-only validation, computational cost assessment, PCA interpretability analysis
Computational
metrics
Training time, inference latency, memory
usage, accuracy degradation
Implementation
platform
Python 3.10, scikit-learn library, XGBoost library
ReproducibilityPublic GitHub repository
Table 5. Baseline supervised classification performance.
Table 5. Baseline supervised classification performance.
ModelAccuracyF1-ScoreBalanced Accuracy
Random Forest0.980.980.98
Decision Tree0.960.960.96
Logistic
Regression
0.950.950.95
Majority
Baseline
0.770.450.33
Table 6. Classification performance under repeated stratified cross-validation.
Table 6. Classification performance under repeated stratified cross-validation.
Feature SpaceModelAccuracyF1-ScoreBalanced Accuracy
OriginalRF0.98 ± 0.010.980.98
OriginalSVM0.97 ± 0.010.970.97
PCARF0.97 ± 0.010.970.97
PCASVM0.95 ± 0.010.950.95
Kernel PCARF0.93 ± 0.020.930.93
Kernel PCASVM0.83 ± 0.020.820.83
Table 7. Feature ablation results.
Table 7. Feature ablation results.
Removed Feature GroupAccuracyF1-ScoreBalanced Accuracy
None0.980.980.98
Distance0.980.980.98
Weather0.970.970.97
Channel metrics0.970.970.97
Physical parameters0.980.980.98
Table 8. LOSO generalization performance.
Table 8. LOSO generalization performance.
Test ScenarioAccuracyF1-ScoreBalanced Accuracy
Clear0.990.990.99
Fog0.970.970.97
Worst0.950.950.95
Rain0.930.930.93
Table 9. Summary cross-scenario transfer performance.
Table 9. Summary cross-scenario transfer performance.
Training ScenariosTest ScenariosAccuracyF1-ScoreBalanced Accuracy
Rain, WorstClear, Fog0.990.990.99
Fog, RainClear, Worst0.980.980.98
Clear, WorstFog, Rain0.970.970.97
Fog, WorstClear, Rain0.950.950.95
Clear, FogRain, Worst0.950.950.95
Clear, RainFog, Worst0.550.540.55
Table 10. Environmental-only validation results.
Table 10. Environmental-only validation results.
ModelAccuracyF1Balanced Accuracy
XGB0.98120.97310.9695
GB0.98100.97360.9724
RF0.98080.97250.9672
DT0.97560.96730.9670
SVM0.94660.91860.9051
MLP0.94660.91620.9043
LR0.89380.81940.8043
Majority0.77100.29020.3333
Table 11. Performance obtained using the original feature space, PCA representations, and Kernel PCA embeddings.
Table 11. Performance obtained using the original feature space, PCA representations, and Kernel PCA embeddings.
ModelOriginalPCAKPCA
RF0.98260.96720.9380
GB0.98240.96740.9314
XGB0.98080.96820.9380
MLP0.97780.95760.8454
DT0.97620.95520.9174
SVM0.97140.95740.8406
LR0.98060.95280.8012
Table 12. Statistical significance analysis comparing the original feature space, PCA-reduced representations, and Kernel PCA embeddings using repeated stratified cross-validation (5 folds × 5 repetitions).
Table 12. Statistical significance analysis comparing the original feature space, PCA-reduced representations, and Kernel PCA embeddings using repeated stratified cross-validation (5 folds × 5 repetitions).
ComparisonMean
Accuracy (Method 1)
Mean
Accuracy (Method 2)
Paired t-Test (p-Value)Wilcoxon
(p-Value)
Cohen’s d
RF Original vs. PCA0.983160.968482.38 × 10−161.2 × 10−54.03
RF Original vs. KPCA0.983160.937084.67 × 10−241.2 × 10−58.64
XGB Original vs. PCA0.981360.969322.52 × 10−151.2 × 10−53.63
XGB Original vs. KPCA0.981360.936762.95 × 10−201.2 × 10−55.96
GB Original vs. PCA0.983000.966361.37 × 10−161.2 × 10−54.13
Table 13. Computational cost associated with the best-performing models.
Table 13. Computational cost associated with the best-performing models.
MethodTrain (s)Infer (ms/Sample)MemoryAccuracyLoss
RF Original1.940.0271.96 MB0.98000
RF PCA1.820.0311.95 MB0.96801.20
RF KPCA2.410.147188.86 MB0.93314.69
XGB Original1.130.0091.94 MB0.98000
XGB PCA0.900.0111.95 MB0.96841.16
XGB KPCA1.510.139188.72 MB0.93164.84
Table 14. PCA-explained variance and physical interpretation of the first principal components.
Table 14. PCA-explained variance and physical interpretation of the first principal components.
Principal
Component
EVRCumulative EVRDominant Physical Interpretation
PC10.32770.3277Global link quality, mainly RF/THz path loss, SNR and capacity
PC20.16280.4905Atmospheric propagation conditions, mainly humidity, visibility and
turbulence
PC30.09400.5846Scenario-dependent variability
PC40.08140.6659Secondary propagation and transition effects
PC50.05950.7254Residual variance associated with
correlated channel metrics
Table 15. Physics-inspired analytical baseline performance.
Table 15. Physics-inspired analytical baseline performance.
RuleAccuracyF1-ScoreBalanced
Accuracy
Min-Eb feasible rule0.98300.97700.9649
Max-SNR rule0.84840.51820.6714
Max-capacity rule0.30500.21160.3987
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Pires, L.M.; Fialho, V. Physics-Supported Linear and Nonlinear Dimensionality Reduction for Supervised Adaptive Channel Selection in Hybrid RF-FSO-THz Communication Systems. Electronics 2026, 15, 2778. https://doi.org/10.3390/electronics15132778

AMA Style

Pires LM, Fialho V. Physics-Supported Linear and Nonlinear Dimensionality Reduction for Supervised Adaptive Channel Selection in Hybrid RF-FSO-THz Communication Systems. Electronics. 2026; 15(13):2778. https://doi.org/10.3390/electronics15132778

Chicago/Turabian Style

Pires, Luis Miguel, and Vitor Fialho. 2026. "Physics-Supported Linear and Nonlinear Dimensionality Reduction for Supervised Adaptive Channel Selection in Hybrid RF-FSO-THz Communication Systems" Electronics 15, no. 13: 2778. https://doi.org/10.3390/electronics15132778

APA Style

Pires, L. M., & Fialho, V. (2026). Physics-Supported Linear and Nonlinear Dimensionality Reduction for Supervised Adaptive Channel Selection in Hybrid RF-FSO-THz Communication Systems. Electronics, 15(13), 2778. https://doi.org/10.3390/electronics15132778

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop