3.1. Instrument-Dependent NIR Spectral Features and Influence of Preprocessing
In
Figure 1, the NIR spectra for benchtop NIR, handheld instrument, and imaging sensor system with and without preprocessing are shown. The spectra from the handheld instrument and the imaging sensor system exhibit a reduced level of structure regarding the number of bands, a consequence of the lower resolution (
Table 1) and overall performance of these instruments. This effect is particularly evident in the imaging sensor system. Its design focuses on robustness and high-throughput industrial applicability rather than detailed chemical resolution, which inherently limits the number of exploitable spectral features and results in a pronounced reduction in spectral detail. The comparatively lower spectral resolution leads to increased band overlap and a reduced ability to resolve weak overtone and combination bands. Therefore, chemical specific information is concentrated in broad absorption regions rather than in distinct spectral features, making these systems more dependent on robust preprocessing strategies.
Consequently, the applied preprocessing methods markedly influenced the spectral appearance (
Figure 1). In particular, the fillPeaks approach reduced low-frequency intensity variations and enhanced the visual prominence of absorption features. Application of the Savitzky–Golay first-derivative transformation resulted in visibly smoother spectral profiles with reduced high-frequency fluctuations, particularly for the handheld and imaging sensor system. This indicates effective attenuation of noise while maintaining the overall band structure [
47]. Beyond the use of the first derivative alone, a combined preprocessing strategy was also evaluated, as suggested by Stipanovic et al. [
33]. Using the combination of the first derivative and SNV increased the intensity of the resulting signals, revealing more pronounced differences between spectra with low and high cotton content.
This observation is illustrated by the bands near 4000 cm
−1 and 5200–5300 cm
−1 (
Figure 1a). In addition, the combination of detrending with SNV appears to be a more promising approach, as evidenced by the observation of distinct differences when compared with the use of detrending or SNV alone. The beneficial effect of SNV- and detrend-based preprocessing might be attributed to the correction of multiplicative scatter effects caused by textile surface roughness, fiber orientation, and path length variations, which are particularly pronounced in fibrous materials.
Characteristic absorption bands for cotton and polyester were assigned based on reported vibrational modes and their corresponding overtone or combination bands in the literature [
14,
22,
33,
48] (indicated by vertical bars in
Figure 1). Weak bands at around 7300 cm
−1 (combination bands of C-H vibrations) [
33,
48] and 7050 cm
−1 (first overtone of O-H stretching modes) [
33,
48] can be seen. A prominent feature of cotton is the band around 6700 cm
−1 that is referenced in literature as first overtone O-H/C-H [
22] with respect to combined O-H stretching and bending vibrations [
14].
The intense band around 5200 cm
−1, assigned to O–H-related overtone and combination modes with respect to overtone of C=O [
22,
48], showed a clear dependence on cotton content [
33]. Additionally, the band at approximately 4700 cm
−1 (attributed to a combination overtone of C–C/C–H and O–H [
22] or to combined O–H stretching and bending vibrations [
14]), and the band at around 4000 cm
−1 (combination of O-H and C-O stretching vibrations [
48]) also appear to be related to the cotton content [
33]. Although several of the discussed absorption bands are associated with O–H vibrations, variations in moisture content can be excluded as a confounding factor, as all samples were conditioned for at least 24 h at ambient temperature and measured within a narrow and controlled time window. Therefore, the observed spectral differences can be attributed primarily to cellulose-related vibrational modes rather than to fluctuations in sample moisture.
In the case of polyester, characteristic absorptions include a band near 6025 cm
−1, which reflects the first overtone of aliphatic and aromatic C–H vibrations [
22,
33]. Additionally, the region around 4425 cm
−1 is associated with combination bands involving C–C, C–H, and C=O vibrations [
22]. A further band appears at around 4100 cm
−1, although no specific assignment was found in the referenced literature.
These spectral characteristics are primarily governed by instrument-specific properties such as spectral resolution and the signal-to-noise ratio, while preprocessing modifies their representation to a lesser extent. They provide the basis for comparing regression models and interpreting differences in predictive performance across the NIR systems investigated.
3.2. Performance Comparison of Regression Models Under Different Preprocessing Conditions
The Savitzky–Golay approach combines local polynomial smoothing with derivative calculation, enabling a reduction in high-frequency noise components when an appropriate window size is selected. As reported in the literature, (excessively) large Savitzky–Golay first-derivative window sizes may lead to oversmoothing and loss of relevant spectral features, whereas very small window sizes can result in insufficient smoothing and increased noise amplification [
47,
49]. Therefore, the influence of different window sizes on model performance was evaluated (
Supplementary Materials, Tables S2–S4).
For partial least squares (PLS) regression, the lowest RMSEP values for both the handheld NIR instrument and the imaging sensor system were obtained using a window size of 11 (handheld: RMSEP = 1.70%; imaging sensor system: RMSEP = 1.15%). For the benchtop NIR system, a window size of 21 yielded the lowest RMSEP (1.12%), closely followed by a window size of 11 (1.16%). A similar trend was observed for the coefficient of determination, with the highest R
2 values achieved at a window size of 11 for the imaging sensor system (R
2 = 0.92) and the handheld instrument (R
2 = 0.81), while the benchtop system reached its highest R
2 at a window size of 21 (R
2 = 0.92). For the kernel-based LS-SVM model, the lowest RMSEP values were consistently obtained using a window size of 11 across all instruments (imaging sensor system: RMSEP = 1.16%; handheld: RMSEP = 1.69%; benchtop: RMSEP = 1.13%). The corresponding R
2 values showed slightly more variation, with the best results achieved at a window size of 5 for the imaging sensor system, 11 for the handheld instrument, and 21 for the benchtop system. However, particularly for the benchtop instrument, the differences between window sizes 11 and 21 were marginal. For the tree-based Random Forest model, the best performance in terms of both RMSEP and R
2 was obtained at a window size of 11 for the imaging sensor system, while a window size of 5 performed best for the handheld instrument. For the benchtop system, similar results were observed for window sizes of 11 and 21, with a slight advantage for the larger window. Overall, these results demonstrate that the choice of the Savitzky–Golay window size is a relevant preprocessing parameter that depends not only on the regression method but also on the spectral resolution and characteristic bandwidths of the respective instrument. To ensure comparability across models and instruments, a window size of 11 was selected for subsequent analyses and visualizations, including the heatmap representation and the combination of the first derivative with SNV normalization, as previously suggested by Stipanovic et al. [
33].
Following the selection of the appropriate Savitzky–Golay first-derivative window size of 11, the influence of different preprocessing strategies on model performance was systematically evaluated (
Figure 2). For PLS regression, the best RMSEP values for the handheld and benchtop NIR instruments were obtained using a combination of detrending and SNV (handheld: RMSEP = 1.62%; benchtop: RMSEP = 1.08%), whereas for the imaging sensor system, the first derivative resulted in the lowest RMSEP (1.15%). R
2 values for the benchtop instrument were largely independent of preprocessing, ranging between 0.91% and 0.92%. For the handheld instrument, the highest R
2 was achieved using the first derivative combined with SNV (0.82), followed by detrending (0.81), while detrend/SNV ranked fourth (0.80). For the imaging sensor system, the first derivative also yielded the highest R
2 (0.92). The relative drop in R
2 between the best and worst preprocessing methods was moderate (approximately 7.5% for the imaging sensor system and handheld, and 3.5% for the benchtop system), whereas RMSEP was much more sensitive, with performance deteriorations exceeding 35%, nearly 20%, and more than 10%, respectively.
For LS-SVM, the optimal preprocessing method depends on the instrument. The lowest RMSEP values were obtained using detrending for the imaging sensor system (RMSEP = 1.04%) and SNV for the handheld (1.57%) with respect to the benchtop instruments (1.09). R2 rankings showed greater variability, with detrending yielding the highest values for the imaging sensor system and handheld instrument, while the first derivative combined with SNV performed best for the benchtop system. The drop from the best to the worst preprocessing method amounted to nearly 20% in RMSEP and over 5% in R2 for the imaging sensor system, around 13% RMSEP and 7% R2 for the handheld instrument, and just over 5% RMSEP and 3% R2 for the benchtop system.
After hyperparameter optimization (see below), Random Forest performance improved overall, but a strong dependence on preprocessing remained. For the imaging sensor system, the first derivative, evaluated across window sizes from 5 to 21, consistently yielded the lowest RMSEP values (1.06–1.19%) and the highest R2 values (0.89–0.91). For the handheld instrument, several preprocessing strategies resulted in comparable RMSEP values (1.58–1.68%), while R2 values remained low and similar (approximately 0.81). For the benchtop system, detrending achieved the lowest RMSEP (0.91) and the highest R2 (0.95), closely followed by the first derivative.
For KRR and XGBoost, similar trends were observed, with the optimal preprocessing strategy being instrument dependent. For both models, differences in R2 across preprocessing methods were generally smaller than those observed for RMSEP. For KRR, the lowest RMSEP values were obtained using detrend/SNV for the benchtop system, SNV for the handheld instrument, and the first derivative for the imaging sensor system, whereas fillPeaks and the combination of first derivative with SNV frequently resulted in elevated prediction errors. For XGBoost, the first derivative consistently achieved the lowest RMSEP for the handheld instrument and imaging sensor system, while detrending performed most favorably for the benchtop system. In both cases, fillPeaks was associated with the lowest predictive accuracy.
In summary, this data suggests that preprocessing is highly data-specific and should not be applied uniformly to all instruments or regression models. The benchtop NIR system, characterized by a high signal-to-noise ratio, exhibited comparatively stable performance across preprocessing methods, whereas handheld and imaging sensor systems showed substantially greater performance variability, indicating a higher sensitivity to preprocessing choices. Methods targeting scatter and baseline effects, such as detrending, SNV transformation, and first derivative-based preprocessing, were particularly important for the handheld instrument and the imaging sensor system. In contrast, fillPeaks exhibited strongly instrument-dependent behavior and frequently degraded performance for high-quality spectra by suppressing fine spectral structures. These findings emphasize that preprocessing decisions are especially critical for lower-resolution or noisier measurement systems, while their impact is reduced for high-quality benchtop instruments. Although the preprocessing strategy influenced performance within each instrument class, absolute differences between instrument types were substantially larger than differences introduced by preprocessing alone. This indicates that preprocessing acts primarily as a modulation factor within instrument-specific signal constraints rather than as an independent determinant of predictive accuracy.
3.3. Comparison of Regression Models and Preprocessing Strategies
Given the high dimensionality and strong multicollinearity of NIR spectroscopic data, overfitting represents a central challenge in model development. Therefore, model evaluation was not based only on absolute predictive performance (RMSEP, R
2), but explicitly on the agreement between training and independent test data. In this context, RMSEC (training data set) and RMSEP (independent test prediction) were evaluated separately. All metrics, together with the corresponding R
2 values for training and test data, are reported to allow for assessment of model generalization and potential overfitting. In addition, model complexity measures (number of latent variables for PLS and selected hyperparameters for machine learning models) were documented to ensure a fair comparison between preprocessing strategies and regression approaches (
Supplementary Materials). During initial model development, tree-based ensemble methods, particularly random forest and XGBoost, exhibited near-perfect fits on the training data for several preprocessing strategies (RMSEC random forest first derivative with SNV benchtop and imaging sensor system <0.05%). This behavior necessitated a more restrictive and carefully controlled hyperparameter optimization to reduce model variance and improve generalization. Therefore, random forest hyperparameters were regularized by limiting tree depth, increasing minimum split and leaf sizes, and enabling bootstrap aggregation (
Supplementary Materials, Table S1), which substantially reduced overfitting and improved the consistency between training and test performance. The other regression models were also trained using model-intrinsic regularization or restricted hyperparameter settings (
Supplementary Materials, Table S1) that corresponded to their respective learning principles. This approach made it possible to focus on model comparisons on generalization behavior rather than on differences resulting from the unrestricted flexibility of the models.
Against this background, PLS regression was used as a linear reference model to provide a robust reference for predictive performance across the different instruments. The optimal number of latent variables varied depending on preprocessing strategy and instrument class and is reported in the GitHub repository. Overall, PLS exhibited stable but comparatively limited predictive performance. For the benchtop NIR system, detrending and detrending combined with SNV preprocessing yielded the best results, with RMSEP values around 1.08–1.11% and R2 values of approximately 0.90–0.92. In contrast, the handheld instrument exhibited substantially higher RMSEP values (approximately 1.62) and greater variability in R2, reflecting the reduced spectral resolution and increased noise level of this system. For the imaging sensor system, first-derivative preprocessing resulted in the best PLS performance (RMSEP 1.15%; R2 0.92), indicating that preprocessing and enhancement in broad spectral trends are particularly important for low-resolution measurements.
LS-SVM models provided improved performance compared to PLS by enabling non-linear regression within a regularized framework. Model performance was strongly dependent on preprocessing strategy, with SNV-based approaches generally yielding the most favorable results. For the benchtop instrument, LS-SVM achieved performance comparable to PLS, whereas for the handheld instrument and the imaging sensor system, clear improvements in RMSEP and R2 were observed. Hyperparameter optimization revealed instrument-dependent regularization requirements, with lower regularization parameters favored for the benchtop system and higher values required for the noisier handheld and imaging sensor system data. These results indicate that LS-SVM can adapt to varying data quality while maintaining stable generalization behavior.
Kernel ridge regression consistently delivered strong generalization performance. Benefiting from its explicit L2 regularization, KRR effectively controlled model complexity in the presence of strong multicollinearity. For the benchtop NIR system, the combination of detrending and SNV preprocessing with KRR resulted in the lowest RMSEP values among all evaluated models, representing a substantial improvement of 10% compared to PLS regression. Similar, though less pronounced, improvements were observed for the handheld instrument. In contrast, first-derivative preprocessing frequently degrades KRR performance, emphasizing the importance of aligning preprocessing strategies with model characteristics. Optimal regularization parameters [
50,
51] were consistently found in the low-to-moderate range (α ≈ 0.01–0.05), corresponding to minima in validation error and indicating a favorable bias–variance trade-off.
Although
XGBoost demonstrated high predictive potential, exemplified by an improvement of up to 9% for the handheld instrument compared to PLS regression, it exhibited a pronounced tendency towards overfitting, particularly for the benchtop and imaging sensor system. Even when the hyperparameter space was constrained to flat or moderately deep trees, reduced learning rates, and subsampling, substantial discrepancies between training and test performance were frequently observed. This behavior can be attributed to the sequential nature of gradient boosting, whereby residuals are iteratively corrected following the framework introduced by Friedman [
52], which can amplify noise and instrument-specific artefacts in high-dimensional and strongly correlated spectral data. Hyperparameter optimization consistently favored strongly regularized configurations with shallow trees and subsampling; nevertheless, predictive performance remained unstable, particularly for the imaging sensor system, highlighting the high sensitivity of boosting methods to data quality and sample size. This behavior is consistent with previous benchmarking studies reporting high tunability but limited robustness of boosting models compared to random forests [
53]. Future work could further improve model robustness by extending the hyperparameter space to include explicit regularization terms such as minimum child weight, split loss thresholds (gamma), and L1/L2 penalties, as well as by applying early stopping strategies based on independent validation sets or nested cross-validation.
The
random forest regression model demonstrated strong and comparatively stable predictive performance across all instruments, with RMSEP values ranging from approximately 0.91–1.19% for the benchtop system, and with moderate performance for the imaging sensor system (RMSEP ≈ 1.19%, R
2 ≈ 0.91). The handheld instrument exhibited slightly higher errors, with RMSEP values ranging from 1.65–1.68%, and with R
2 values ranging from 0.81 to 0.82. By constraining tree depth, enforcing conservative node splitting criteria, and enabling bootstrap aggregation, model variance was effectively controlled, resulting in improved agreement between training and test performance. Inspection of the optimized random forest hyperparameters revealed a high degree of consistency across instruments. Optimal configurations were characterized by moderate tree depths (max_depth = 10–20), conservative splitting criteria, and standard feature subsampling strategies (√ or log2), indicating that performance gains were achieved without excessive model complexity, consistent with the inherent variance-reducing design of random forests [
54]. Notably, the optimization procedure consistently selected the lowest allowed value for the minimum number of samples per leaf (min_samples_leaf = 5) across instruments, indicating a tendency of the model to favor highly flexible tree structures. Due to the rapidly increasing computational cost associated with larger hyperparameter spaces, the range of evaluated random forest configurations had to be deliberately constrained. The discussed observations are consistent with the findings of previous large-scale benchmark studies, which have reported comparatively low tunability of random forest models and a tendency towards robust, near-default configurations [
50,
53].
Summarized for the benchtop NIR system, advanced regression models consistently achieved high predictive accuracy, with RMSEP values around or below 1.00% and R2 values of approximately 0.94–0.95 for RF, XGBoost, LS-SVM, and KRR. These results clearly outperform standard PLS regression, which yielded an RMSEP of 1.08% and an R2 of 0.91. The high signal-to-noise ratio and spectral resolution of the benchtop instrument enabled robust exploitation of nonlinear relationships, particularly when combined with appropriate preprocessing.
For the handheld NIR instrument, model performance was more heterogeneous and less consistent. Although advanced models generally outperformed PLS regression, prediction errors were higher overall. RMSEP values between 1.46 and 1.56% with corresponding R2 values of 0.81–0.85 were observed, compared to PLS regression with an RMSEP of 1.62% and an R2 of 0.80. Notably, random forest did not consistently outperform PLS for the handheld system, indicating that increased model complexity does not necessarily translate into improved performance for lower-resolution or noisier instruments. For the imaging sensor system, gradient boosting models frequently exhibited unstable behavior and pronounced overfitting. The best XGBoost results reached an RMSEP of approximately 1.25% with an R2 of 0.88, whereas the remaining regression models achieved substantially more robust performance, with RMSEP values between 1.03 and 1.13% and R2 values of 0.90–0.93. These results are comparable to or slightly better than those obtained with PLS regression (RMSEP = 1.15%, R2 = 0.92) and exceed the predictive performance observed for the handheld instrument, despite the lower spectral resolution of the imaging sensor system.
Across all instruments, no single regression model or preprocessing strategy was universally optimal. In line with the No Free Lunch Theorem [
41], model performance was strongly context-dependent and governed by instrument-specific characteristics such as spectral resolution, noise level, wavelength coverage, and the interaction between preprocessing and model-intrinsic regularization. The comparison further demonstrates that preprocessing strategies not only affected predictive accuracy but also influenced model complexity, reflected in varying numbers of latent variables for PLS and instrument-dependent hyperparameter configurations for nonlinear models. Under more heterogeneous real-world conditions, these interactions are expected to become even more pronounced. Increased surface roughness, dyeing, and structural variability would likely enhance the relevance of scatter-correction methods such as SNV or MSC, while derivative-based preprocessing may help suppress broad background contributions but simultaneously amplify noise in lower-resolution inline systems. Spectral regions dominated by cellulose-specific O–H combination bands may remain comparatively robust, whereas weaker overtone regions could lose discriminative power due to overlapping contributions from colorants and additives.
3.4. Model Interpretability: Coefficients and Feature Relevance
Feature importance analysis was performed for XGBoost and partial least squares (PLS) regression, as both provide interpretable measures of variable relevance. In XGBoost, importance is derived from variable contributions to tree splits [
55], while PLS relates spectral variables to latent components maximizing covariance with the response, allowing for physically meaningful interpretation [
34,
56].
XGBoost Model Interpretability
The XGBoost feature importance analysis (
Figure 3a) was performed separately for the imaging sensor system, the handheld device, and the benchtop system. For each instrument, the preprocessing strategy yielding the most robust generalization performance was selected (imaging sensor system: fillPeaks; handheld: first derivative resp. SNV; benchtop: detrend + SNV). The interpretation therefore focuses on system-specific, model-relevant features rather than on preprocessing-independent trends.
For the imaging sensor system, the most important XGBoost features are predominantly located in the region between approximately 7600 and 7700 cm−1, with additional contributions above 7000 cm−1. These wavenumbers dominate the model decision and reflect the limited spectral resolution and overall performance of the imaging sensor system, which favors statistically stable overtone regions over well-resolved band structures. Among the identified features, a contribution around 7300 cm−1, attributed to combination bands of C–H vibrations, overlaps with a literature-reported cotton-related absorption. Other characteristic cotton bands, such as those around 6730 or 5200 cm−1, are not relevant for this system and do not appear in the feature importance ranking. Polyester-related features are likewise not dominant in the imaging sensor system data.
For the
handheld instrument, XGBoost feature importance was evaluated using the two best preprocessing strategies, first derivative and SNV, both of which yielded comparable predictive performance. A direct comparison of the feature rankings obtained with these two preprocessing approaches reveals a high degree of consistency in the spectral regions exploited by the model. In both cases, the most important features cluster around the cotton-related bands near 5200 cm
−1 and 4700 cm
−1. Although the exact numerical positions and ranking order of individual wavenumbers differ between first derivative and SNV preprocessing, these differences represent small shifts within broad near-infrared bands. In contrast, polyester-related absorptions around 4400 cm
−1, which are clearly visible in the preprocessed NIR spectra (
Figure 1), do not appear among the most important XGBoost features. This observation indicates that, while these bands are spectrally present, they do not contribute significantly to the statistical separation achieved by the model. Instead, the XGBoost classifier primarily relies on cotton-associated spectral regions. This behavior reflects both the limited spectral resolution of the handheld instrument and the tendency of tree-based models to select the most statistically robust spectral regions rather than all spectroscopically observable features.
For the benchtop system, which achieved the highest predictive performance, the most important XGBoost features are concentrated at discrete wavenumbers, most notably around 6260 cm−1, with additional relevant contributions in the regions around 5000 cm−1 and 4000 cm−1. The high signal-to-noise ratio and spectral resolution of the benchtop system enable the model to rely on a small number of robust and reproducible spectral features. With respect to material-specific bands, only a limited subset of previously assigned absorptions contributes to the model decision. Cotton-related features are observed at ~7300 cm−1 and ~4069 cm−1, whereas other literature-reported cotton bands (e.g., around 7050, 6730, 5200, or 4700 cm−1) do not appear among the most important XGBoost features. For polyester, contributions at around 4100 cm−1 and 6025 cm−1 appear in the features, while other polyester-related bands are not relevant in the final model.
Across all systems, XGBoost selects point-wise spectral features that provide the highest statistical discrimination for the respective instrument and preprocessing strategy. While not all literature-reported cotton and polyester bands contribute equally to the models, several of the features coincide with known NIR overtone and combination bands. Therefore, the selected features should not be interpreted as isolated chemically specific signals but rather as representative points within spectrally relevant regions. In this context, PLS regression provides a more stable and physically interpretable representation of band structures.
PLS Model Interpretability
For PLS regression, latent variables 2–4 were interpreted, as the first latent variable mainly captured dominant, globally correlated spectral variance rather than material-discriminative chemical information.
For the
imaging sensor system, PLS loadings of components 2–4 are broad and weak, reflecting the limited spectral resolution and restricted wavelength coverage of the system (
Figure 3b). Alternating positive and negative loadings appear around 9000–8800 cm
−1 and near 8600 cm
−1, but these features cannot be unambiguously linked to specific molecular vibrations. The interval between 8400 and 8000 cm
−1 is largely featureless, with only weak contributions near 7800 cm
−1 and slightly more pronounced but unspecific features around 7600 cm
−1. Chemically more meaningful contributions emerge at lower wavenumbers. Positive loadings for components 2 and 4 are observed around 7300 and 6700 cm
−1, coinciding with literature-reported cotton-related overtone and combination bands. A feature near 6400 cm
−1 does not correspond to known cotton or polyester absorptions and is interpreted as non-specific variance.
For the handheld NIR system, PLS loadings of components 2–4 exhibit more pronounced and structured features. Positive loadings around 6000 cm−1 can be associated with polyester-related bands, indicating that the handheld instrument captures material-specific information from both fiber components. Strong negative loadings occur around 5600 cm−1, although no clear cotton or polyester assignment could be identified, suggesting overlapping combination bands or systematic spectral variance. Negative loadings near 5100 cm−1 are observed close to the well-established cotton-related band at around 5200 cm−1, indicating contributions from cotton-associated bands. Additional positive and negative loadings below 4400 cm−1 overlap with cotton-related absorptions in the 4400–4500 cm−1 region. Particularly intense features for components 2 and 3 are found between 4800 and 4900 cm−1, a region frequently associated with cotton-dependent combination bands. Overall, both cotton- and polyester-related regions contribute to the latent variables, although several strong features remain difficult to assign unambiguously.
The benchtop NIR system shows the most distinct and chemically interpretable PLS loadings, consistent with its superior spectral resolution, broader wavelength range, and highest predictive performance. Negative loadings for components 2–4 are observed around 7000–7100 cm−1, coinciding with cotton-related overtone bands. Strong positive and negative loadings occur between 6200 and 6100 cm−1, associated with polyester-related overtones, as well as near 5200–5300 cm−1, corresponding to prominent cotton-related bands. Additional intense features are present around 4700 cm−1, which are well-known to be sensitive to cotton content, while smaller features near 4400 cm−1 coincide with polyester-related absorptions. Below 4200 cm−1, several features of varying intensity overlap with both cotton- and polyester-related bands, highlighting the ability of the benchtop system to capture chemically meaningful variance across multiple relevant spectral regions.
Across all instruments, cotton-related spectral regions contribute to the latent variables used for prediction; however, their number, clarity, and interpretability strongly depend on the instrument class. The imaging sensor system captures only a small subset of cotton-associated bands, the handheld system reflects contributions from both cotton and polyester with several ambiguous features, and the benchtop NIR system exhibits the most comprehensive and chemically interpretable loading structure. The clearer chemical interpretability observed for the benchtop system further supports the conclusion that instrument-dependent signal quality, rather than preprocessing alone, governs the robustness of spectroscopic modeling.
While XGBoost identifies point-wise spectral positions that maximize statistical discrimination within non-linear decision rules, PLS loadings represent distributed spectral patterns across broader absorption regions. An alternative to interpreting individual loadings would be the calculation of variable importance in projection (VIP) scores [
57], which provide a concise summary of the overall relevance of each variable. However, in the present study, loadings were intentionally examined to preserve the structure of the latent variables. As XGBoost feature importance is driven by statistical contributions to model decisions, spectral interpretation was mainly guided by PLS regression, which provides physically meaningful loadings linked to spectral covariance with the response variable. More advanced model-agnostic interpretability approaches such as SHAP (SHapley Additive exPlanations) values [
58] may provide additional insights into non-linear model behavior, but these were beyond the scope of the present study.
3.5. Model Training Time
The influence of spectral pre-treatment on training time varied substantially across models and instruments (
Supplementary Materials, Tables S2–S4). It should be noted that the reported training times also depend on the hyperparameter optimization budget, i.e., the size of the search grid. Consequently, absolute runtimes are not solely a property of the learning algorithm and preprocessing, but scale approximately with the number of evaluated parameter combinations. For PLS regression, preprocessing had only a minor effect on computational cost. Neither detrending, SNV, nor different Savitzky–Golay window sizes altered training time meaningfully for the handheld or imaging sensor system instruments. The benchtop system exhibited a slightly higher variability (coefficient of variation: 15.9%), yet, even here, the differences between pre-treatments remained moderate. Among the pre-treatments, fillPeaks resulted in the longest PLS training times, whereas window-size effects of Savitzky–Golay filters remained negligible. Across instruments, the handheld NIR exhibited 1.37-fold longer PLS training times than the imaging sensor system, whereas the benchtop instrument required 351-fold longer training times. The generally higher training times for benchtop PLS models are primarily attributable to the substantially higher spectral resolution, which increases matrix dimensionality and thereby amplifies the scaling behavior of PLS. This reflects the dominant role of spectral dimensionality rather than the chosen pre-treatment. KRR displayed a highly consistent computational profile across pre-treatments, as indicated by low coefficients of variation (handheld: 4.93%, imaging sensor system: 6.75%, benchtop: 7.39%). Relative to the handheld instrument, the imaging sensor system was approximately 2.4 times faster, whereas benchtop data resulted in a 2.6 times slower training process. A similar pattern was observed for LS-SVM. Here, pre-treatments again had negligible influence, including the Savitzky–Golay window size. The variability remained low to moderate across instruments (handheld: 6.69%, imaging sensor system: 17.7%, benchtop: 1.26%). Imaging sensor system data enabled the fastest training (4 times faster than handheld), whereas benchtop training times were modestly slower (1.7 times) than for handheld. Tree-based models showed the largest pre-treatment-related variability. For random forest, differences between window sizes and preprocessing strategies were more pronounced, reflected by substantially higher coefficients of variation (handheld: 31.1%, imaging sensor system: 30.9%, benchtop: 27.5%) compared to the other regression approaches. Despite this variability, the ranking of devices was consistent: the imaging sensor system enabled the fastest training (2.5 times faster than handheld), while benchtop spectra led to nearly tenfold higher training times. XGBoost displayed a similar behavior. Although differences between pre-treatments were visible (coefficient of variation: 13–16%), no systematic trend emerged. Again, the imaging sensor system produced the lowest training times (1.5 times faster than handheld), whereas benchtop spectra resulted in a nearly tenfold increase (9.5 times slower). Overall, the influence of preprocessing on training time was strongly model-dependent. For PLS and the kernel methods, pre-treatments introduced only minor runtime variation, indicating algorithmic robustness to spectral transformations. In contrast, tree-based ensembles were more sensitive to pre-treatment, likely due to their dependence on local variance structure and feature distributions. Savitzky–Golay filtering, especially with a window size of 21, often yielded the fastest runtimes across all instruments, although this pattern was not consistent.
Building upon these pre-treatment-specific observations, the relative training times across models and instruments were compared to evaluate the overall computational behavior of the applied regression methods (
Figure 4). PLS was used as the reference since this is the most common method in the literature. Across the handheld NIR and imaging sensor system, kernel-based methods (LS-SVM and KRR) require approximately one order of magnitude more training time than PLS, whereas ensemble models (random forest and XGBoost) exhibit increases of two to three orders of magnitude. These patterns reflect the higher algorithmic complexity of tree-based ensembles in high-dimensional spectral data, where repeated split evaluation and large model sizes dominate computation. The relatively small standard deviations for LS-SVM and KRR indicate stable computational behavior across preprocessing variants. By contrast, the greater variability observed for random forest is consistent with the sensitivity to preprocessing-induced changes. The computational costs for the handheld seem to be somewhat higher than with the imaging sensor system. In contrast, the benchtop dataset shows a different profile. Here, LS-SVM and KRR train substantially faster than PLS (ratios < 0.1), while the ensemble models remain more expensive but at markedly lower relative levels than for the other instruments. This is consistent with the previously mentioned explanation, as the higher spectral resolution increases the number of variables and thereby amplifies the computational cost of PLS regression.
Overall, the results demonstrate that PLS remains the most computationally efficient model across all instruments. Kernel methods result in moderate additional cost, but their efficiency can strongly depend on spectral quality and instrument characteristics. Tree-based ensemble models, as expected, are the most computationally intensive, particularly for lower-quality or more variable spectra. The reported computational costs refer exclusively to model training during method development and do not reflect real-time prediction, which is substantially faster once the model is deployed.