Towards Early-Stage Corrosion Prediction Using UHF RFID Measurements: A Machine Learning Feasibility Study
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis manuscript proposes an intelligent corrosion monitoring framework that combines UHF RFID tag-antenna sensing with machine learning. The authors use RFID-derived features, including AID, forward power, frequency, RSSI/backscattered power, and phase, to perform early-stage corrosion classification and nominal corrosion thickness estimation. May have certain engineering useful. However, the current manuscript has significant weaknesses in experimental rigor, data independence, validation design, task definition, and interpretation of the reported results.
- Insufficient data independence and limited generalization.
The manuscript reports 1089 RFID records from four corrosion stages. However, these records should not be treated as 1089 independent experimental samples. Many of them appear to be repeated readings or frequency-sweQep records acquired from the same physical specimens under the same experimental conditions. If records from the same specimen, corrosion stage, readcount session, or frequency sweep are split across training and test sets, the model may learn dataset-specific patterns rather than corrosion-related features that generalize to new samples. The authors introduced Leave-One-Sample-Out validation, but the definition of “sample” seems to refer to readcount or measurement session rather than an independent physical corrosion specimen. This does not demonstrate generalization to new steel plates, new corrosion morphologies, new tag placements, new reader-tag distances, or new environmental conditions. The authors should perform group-wise validation based on truly independent physical specimens. If only one physical specimen exists for each corrosion stage, the manuscript should not claim robust generalization.
- 2. AID formulation and geometry-cancellation claim are too strong. In realistic UHF RFID measurements, multipath, polarization mismatch, tag orientation, metallic boundary effects, chip nonlinearity, and reader dynamics may prevent exact cancellation. The statement should be softened to “partially compensates” or “reduces sensitivity under ideal assumptions,” unless the authors can experimentally demonstrate exact cancellation.
- 3. The manuscript reports ARI = 1.0, NMI = 1.0, R² = 1.00, and LOSO RMSE = 0.12 μm. Such near-perfect results are unusual for real corrosion monitoring and require careful explanation. Additional analyses are needed, including confusion matrices, residual plots, per-stage error distributions, confidence intervals, bootstrap analysis, random-seed stability, feature ablation, and comparisons with simple rule-based or threshold-based baselines. If a simple AID or forward-power threshold achieves similar performance, the necessity of more complex machine learning models should be reconsidered.
- 4. The manuscript is written as if it validates a complete intelligent RFID corrosion monitoring framework. However, the dataset appears to be inherited from a prior experimental study. If the contribution of the present manuscript is mainly machine-learning analysis of an existing dataset, this should be explicitly stated in the title, abstract, and contribution section. The conclusions should be moderated accordingly.The authors are encouraged to include an independent test dataset involving new steel specimens, new corrosion exposure durations, different tag placements, different reader-tag distances, or different environmental conditions. Without such validation, the manuscript should be framed as a feasibility study rather than a validated monitoring system.
Author Response
Comment 1: Insufficient data independence and limited generalisation. The manuscript reports 1089 RFID records from four corrosion stages. However, these records should not be treated as 1089 independent experimental samples. Many of them appear to be repeated readings or frequency-sweep records acquired from the same physical specimens under the same experimental conditions. If records from the same specimen, corrosion stage, readcount session, or frequency sweep are split across training and test sets, the model may learn dataset-specific patterns rather than corrosion-related features that generalise to new samples. The authors introduced Leave- One-Sample-Out validation, but the definition of “sample” seems to refer to readcount or measurement session rather than an independent physical corrosion specimen. This does not demonstrate generalisation to new steel plates, new corrosion morphologies, new tag placements, new reader-tag distances, or new environmental conditions. The authors should perform group-wise validation based on truly independent physical specimens. If only one physical specimen exists for each corrosion stage, the manuscript should not claim robust generalisation.
How we addressed:
We thank the reviewer for this precise and important observation and agree entirely. The 1,089 records do not represent 1,089 independent experimental samples. This has been explicitly acknowledged in the revised manuscript in Section 3.1 (lines 195--203), which now states that the 1,089 entries represent RFID measurement records rather than independent physical corrosion specimens, that each corrosion-stage file contains multiple RFID reads collected during frequency-sweep measurements repeated several times for the same underlying specimen, and that the dataset therefore comprises repeated observations of a limited number of physical corrosion samples rather than a collection of fully independent experimental specimens. Multiple records originate from the same physical specimen under comparable experimental conditions, differing primarily in operating frequency, reader activation conditions, and repeated measurement events.
The LOSO validation groups records by their cumulative readcount value, providing separation at the measurement-event level to prevent within-session leakage, but does not constitute specimen-level independence since all records originate from the same four physical steel plates. The term "sample" in the LOSO procedure refers to a measurement-event grouping based on the readcount variable rather than an independent physical corrosion specimen, and LOSO validation in this study evaluates generalisation across repeated measurement events while reducing within-session data leakage. This is now explicitly stated in Section 4.2.2 (lines 488--491). All language claiming robust generalisation beyond the tested conditions has been removed from the manuscript.
Since the dataset contains only one physical specimen per corrosion stage by experimental design a constraint inherited from the original Zhang et al. [7] study group-wise validation based on truly independent physical specimens cannot be performed with the available data. This is acknowledged in the limitations section (lines 618--626), which states that the LOSO validation strategy does not establish generalisation to new physical steel plates, different corrosion morphologies, alternative RFID tag placements, different reader-tag geometries, or varying environmental conditions, and that since the dataset originates from a single physical plate for each corrosion stage, the model may partially capture specimen-specific electromagnetic characteristics in addition to corrosion-related behaviour. Validation using independent steel plates, additional corrosion morphologies, and measurements collected under varying operating conditions is therefore required before strong conclusions regarding deployment-level generalisation can be made. Collecting data from multiple independent specimens across varied conditions is identified as the primary direction for future work.
Comment 2: AID formulation and geometry-cancellation claim are too strong. In realistic UHF RFID measurements, multipath, polarisation mismatch, tag orientation, metallic boundary effects, chip nonlinearity, and reader dynamics may prevent exact cancellation. The statement should be softened to “partially compensates” or “reduces sensitivity under ideal assumptions,” unless the authors can experimentally demonstrate exact cancellation.
How we addressed:
We thank the reviewer for this precise observation and agree entirely. The original manuscript overstated the geometric compensation provided by the AID formulation by implying exact cancellation of orientation and distance dependencies. The revised manuscript has addressed this in three locations.
In the Literature Review (lines 88--91), the description of AID has been revised from "constructed to cancel orientation and distance dependencies" to "a dimensionless power ratio designed to reduce sensitivity to orientation and distance dependencies under ideal measurement conditions, isolating the antenna impedance change induced by corrosion as the primary measurable quantity," ensuring appropriately qualified language from the first introduction of AID.
In Section 3.2 (lines 245--252), the language has been further revised to explicitly state that the geometric compensation provided by AID is only approximate and is derived under idealised line-of-sight propagation assumptions. In practical corrosion-monitoring environments, additional electromagnetic effects, including multipath propagation, polarisation mismatch, metallic edge diffraction, chip nonlinearity, multiple rust layers, and non-uniform oxide growth, may introduce coupling losses that are not fully cancelled within the AID formulation. Consequently, AID should be interpreted as a feature with reduced sensitivity to reader-tag geometry rather than complete immunity to geometric and environmental influences. The phrase "cancel exactly" has been replaced throughout the manuscript with "partially compensate" and "reducing the sensitivity of AID to reader-tag geometry under ideal measurement conditions."
In Section 5 (lines 633--639), the limitations paragraph now explicitly acknowledges that the AID formulation is derived under idealised propagation assumptions and therefore provides partial rather than absolute compensation of reader-tag orientation and distance effects, and that practical factors such as edge diffraction, multipath propagation, polarisation mismatch, oxide-layer inhomogeneity, localised pitting corrosion, and multilayer rust growth may introduce additional coupling effects that are not explicitly modelled.
Comment 3: The manuscript reports ARI = 1.0, NMI = 1.0, R² = 1.00, and LOSO RMSE = 0.12 µm. Such near- perfect results are unusual for real corrosion monitoring and require careful explanation. Additional analyses are needed, including confusion matrices, residual plots, per-stage error distributions, confidence intervals, bootstrap analysis, random-seed stability, feature ablation, and comparisons with simple rule-based or threshold-based baselines. If a simple AID or forward-power threshold achieves similar performance, the necessity of more complex machine learning models should be reconsidered.
How we addressed:
Explanation of near-perfect results. The structural explanation of the near-perfect LOSO results is provided in detail in Section 4.2.2 (lines 496--500), where the fold construction table (Table 6) shows that six of the seven folds contain zero M6 records in their test sets, yielding RMSE of 0.000 µm by construction. Fold 6 is the only fold containing all four corrosion stages and yields a Random Forest RMSE of 0.817 µm, which is the most meaningful validation result. The near-perfect aggregate RMSE of 0.12 µm is therefore a structural consequence of the readcount-based fold distribution rather than an indication of unrealistic model performance.
Confidence intervals. Confidence intervals are reported throughout Tables 5, 7, and 8 as mean ± standard deviation across all validation folds, providing a direct measure of result stability across different data splits.
Feature ablation. The SHAP analysis presented in Section 4.3 (lines 516--567) provides principled feature importance analysis, quantifying the individual contribution of each feature to model predictions and providing interpretable feature importance information directly linked to the electromagnetic physics of the sensing system.
Confusion matrices. The regression model in this study predicts continuous numerical thickness values rather than discrete class labels. Confusion matrices are an evaluation metric for classification tasks and are therefore not applicable to the regression framework employed in this study.
Bootstrap analysis. Since the dataset contains only four physical steel plates one per corrosion stage bootstrap resampling would draw repeated samples from the same physical specimens and would not provide meaningful additional evidence of generalisation beyond what is already reported through the LOSO validation and fold construction analysis.
Overstated claims softened. Strong claim language including "successfully," "perfect discrimination," and "confirming generalisation" has been removed throughout the revised manuscript. The DBSCAN results section has been revised at lines 415--417 to state that DBSCAN identified four distinct clusters that aligned with the four nominal corrosion stages within the analysed dataset, suggesting that the corrosion stages present in the available dataset exhibit strong separability within the selected feature space. All performance claims are now appropriately bounded by the constraints of the available dataset and validation design.
Comment 4: The manuscript is written as if it validates a complete intelligent RFID corrosion monitoring framework. However, the dataset appears to be inherited from a prior experimental study. If the contribution of the present manuscript is mainly machine-learning analysis of an existing dataset, this should be explicitly stated in the title, abstract, and contribution section. The conclusions should be moderated accordingly. The authors are encouraged to include an independent test dataset involving new steel specimens, new corrosion exposure durations, different tag placements, different reader-tag distances, or different environmental conditions. Without such validation, the manuscript should be framed as a feasibility study rather than a validated monitoring system.
How we addressed:
We thank the reviewer for this important recommendation and agree entirely. The revised manuscript has been reframed as a machine learning feasibility study rather than a validated monitoring system, with changes made in three locations.
The title has been revised to "A Machine Learning Feasibility Study for Early-Stage Corrosion Detection Using UHF RFID Measurements", explicitly positioning the work as a feasibility study from the outset.
The abstract (lines 3--6) now states that "this study presents a machine learning feasibility study for early-stage corrosion detection" and that "machine learning algorithms were applied to a previously published RFID corrosion dataset," making the secondary dataset nature and feasibility framing explicit from the first sentences. The concluding statement of the abstract has been moderated to state that "the results demonstrate the potential of RFID-enabled machine learning for early-stage corrosion assessment and provide a foundation for future experimental validation," replacing any claim of a validated monitoring system.
The conclusions and limitations sections have been moderated accordingly throughout the revised manuscript, with all claims of robust generalisation removed and the need for future experimental validation on independent specimens explicitly identified as the primary next step.
Reviewer 2 Report
Comments and Suggestions for AuthorsThis work proposes an intelligent monitoring framework for early metal corrosion that integrates ultra-high frequency RFID sensing and machine learning. Collect RFID data of corroded steel from 0/1/3/6 months, and construct features such as AID and forward power; Unsupervised DBSCAN clustering can perfectly distinguish four types of corrosion stages without labels, and the effect is much better than K-Means. SHAP analysis confirms that AID and forward power are core sensing features. This passive low-cost solution compensates for the limitations of traditional non-destructive testing, but is only based on laboratory data, and further expansion of complex operating conditions and deep learning models is needed. Before being considered for acceptance, this work should provide explanations or clarifications regarding the following issues, as follows:
[1] The original caption text is mixed with formulas and abbreviations, and there is no unified format. The title of the icon is placed below the image but lacks standard journal format specifications. Some charts (Figure 2, Figure 5) and subgraphs (a) and (b) are labeled too small and do not have independent subgraph explanations; The coordinate axis of the PCA clustering graph is only labeled with principal component 1, without principal component 2 identification, making it difficult to read the graph. The K-Means and DBSCAN clustering maps use colors to distinguish clusters and labels to distinguish the actual corrosion stages, without clear legend explanations, making it easy for readers to confuse. Figure 2 shows that the coordinate axis scales of multiple subgraphs have varying densities, and the curves have no distinguishing markings. The overlapping lines of different corrosion stages are difficult to distinguish.
[2] Table 2AID shows that the rows and columns of the step-by-step calculation table are misaligned and the segments are chaotic, and the readability of the formulas and text is poor. All tables do not have a three line table format, with redundant vertical and horizontal lines, which does not comply with the formatting conventions of journals. The scatter plot for thickness prediction lacks error labeling, making it difficult to visually display the prediction deviation of 43 μ m samples.
[3] In the text, only a single measurement batch is excluded for LOSO, and the physical specimens are still from the same batch of steel plates. Only multiple measurements are repeated, and there is no completely independent new corrosion specimen. The model learns the inherent electromagnetic base characteristics of the same steel plate, rather than pure corrosion characteristics. The ultra-high accuracy may have false optimism and cannot prove that the model can generalize to completely new unknown corroded steel plates.
[4] Corrosion variables are single and mixed, lacking controlled variable control experiments. There is no control group set up, no long-term environmental control of blank non corroded steel plates, and no experiment on the interference of humidity/temperature on RFID signals, making it impossible to distinguish whether signal changes are caused by corrosion of the oxide layer or environmental noise.
[5] Experimental sweep frequency is 902-928MHz, and corrosion changes the resonant frequency of the antenna. Theoretically, the resonant shift varies with different degrees of corrosion, and the frequency itself should have discriminative ability. However, SHAP's frequency contribution of 0 only indicates the limitations of the fixed frequency scanning process design in this experiment, which does not have universality and cannot be extended to variable frequency and multi resonant RFID tags. The conclusion is to use a special case instead of a general rule.
[6] Laser thickness measurement and RFID RF acquisition both have inherent instrument noise, which theoretically cannot be fully fitted; The ultra-high accuracy comes from the fact that the dataset only has 4 discrete thickness labels, which belongs to limited classification fitting rather than continuous regression. The model is essentially a classification disguised regression, and the ability to predict true continuous thickness is overestimated.
[7] The derivation of the AID formula assumes that the reader tag distance and antenna polarization are completely cancelled out, but in reality, the edges of metal components, multiple layers of rust, and non-uniform oxide layers will introduce additional coupling losses. The paper does not discuss the failure boundary conditions of AID and excessively claims that AID is not affected by geometric posture. Only explaining the positive correlation between forward power, AID, and corrosion thickness, without explaining the characteristic nonlinear fluctuations caused by the dielectric constant, roughness, and uneven corrosion of the oxide layer, cannot explain the source of prediction error for M1 (43 μ m), and the combination of mechanism and model results is incomplete.
Author Response
Comment 1: The original caption text is mixed with formulas and abbreviations, and there is no unified format. The title of the icon is placed below the image but lacks standard journal format specifications. Some charts (Figure 2, Figure 5) and subgraphs (a) and (b) are labelled too small and do not have independent subgraph explanations; The coordinate axis of the PCA clustering graph is only labelled with principal component 1, without principal component 2 identification, making it difficult to read the graph. The K-Means and DBSCAN clustering maps use colours to distinguish clusters and labels to distinguish the actual corrosion stages, without clear legend explanations, making it easy for readers to confuse. Figure 2 shows that the coordinate axis scales of multiple subgraphs have varying densities, and the curves have no distinguishing markings. The overlapping lines of different corrosion stages are difficult to distinguish.
How we addressed:
We thank the reviewer for the detailed and constructive observations on figure quality and formatting. Each point has been addressed as follows, with changes.
“The original caption text is mixed with formulas and abbreviations, and there is no unified format”
Ans: All figure captions have been revised to follow MDPI journal formatting standards. Each caption is now a complete, self-contained sentence ending with a full stop, with all abbreviations expanded on first appearance within the caption. AID is expanded as Analogue Identifier, PCA as Principal Component Analysis, DBSCAN as Density-Based Spatial Clustering of Applications with Noise, and LOSO as Leave-One-Sample-Out within their respective captions.
“Some charts (Figure 2, Figure 5) and subgraphs (a) and (b) are labelled too small and do not have independent subgraph explanations”
Ans: Figures 2 and 5, which contain multiple panels, have been reformatted using the MDPI wide figure environment with subfloat so that each panel now carries its own independent subcaption describing its specific content. Readers can understand each panel without referring to the main caption or the body text.
“The coordinate axis of the PCA clustering graph is only labelled with principal component 1, without principal component 2 identification, making it difficult to read the graph”
Ans: The x-axis label in Figure 2 previously read "Frequency (kHz)" while the tick values correctly displayed MHz values. This has been corrected to "Frequency (MHz)" across all four subplots.
“The coordinate axis of the PCA clustering graph is only labelled with principal component 1, without principal component 2 identification, making it difficult to read the graph”
Ans: We respectfully note that both the K-Means and DBSCAN clustering figures already include labels on both axes, with Principal Component 1 on the x-axis and Principal Component 2 on the y-axis. However, the font size of the axis labels has been increased in the revised figures to improve readability.
“The K-Means and DBSCAN clustering maps use colours to distinguish clusters and labels to distinguish the actual corrosion stages, without clear legend explanations, making it easy for readers to confuse.”
Ans: We respectfully note that both figures already contain two clearly separated legends: one distinguishing algorithm-generated clusters by colour, and one distinguishing the true corrosion stages by marker shape. The captions also explicitly state that colours denote algorithm-generated clusters and marker shapes indicate the true corrosion stages. Nevertheless, the legend titles and font sizes have been increased in the revised figures to further improve clarity.
“The overlapping lines of different corrosion stages are difficult to distinguish.”
Ans: The four subplots of Figure 2 have been separated into four individual subfigures, each with a consistent x-axis scale displaying frequency in MHz and an appropriately scaled y-axis. The confidence interval shading provides visual distinction between overlapping corrosion stage curves. To complement the revised figures, a detailed quantitative discussion of each feature's behaviour across corrosion stages has been added to Section 4.1 (lines 373--391). This addition explains the physical basis of the observed feature behaviour: forward power demonstrates the strongest discriminatory capability, increasing substantially from M0 (19--19.5~dBm) to M6 (approximately 25.5~dBm), while AID inversely correlates with corrosion progression, with M0 exhibiting the highest values (0.6--0.8) and M6 the lowest (0.2--0.3). Backscattered power and phase show significant overlap and high variance across stages, consistent with their exclusion as primary features. An explanation of the nonlinear electromagnetic interactions underlying these observations is also provided at lines 392396.
Comment 2: Table 2AID shows that the rows and columns of the step-by-step calculation table are misaligned, the segments are chaotic, and the readability of the formulas and text is poor. All tables do not have a three-line table format, with redundant vertical and horizontal lines, which does not comply with the formatting conventions of journals. The scaler plot for thickness prediction lacks error labelling, making it difficult to visually display the prediction deviation of 43 µm samples.
How we addressed:
All tables have been converted to the standard three-line format, with all redundant inter-row lines removed. Column widths and row spacing in Table 2 have been adjusted to resolve the misalignment and improve the readability of the embedded formulae.
We thank the reviewer for this observation. The scatter plot format itself serves as the visual representation of prediction error deviations from the identity line, directly encode prediction accuracy for each sample. As visible in Figure 5(a), the prediction deviations at the 43 µm stage (M1) are already clearly apparent, with predicted values scattered above the identity line while actual values remain at 43 µm. The physical origin of this deviation has been addressed in the revised Section 4.2.1 (lines 444--453), where a detailed explanation attributes the greater prediction uncertainty at the M1 stage to the spatially non-uniform, thin oxide layer characteristic of early-stage pitting corrosion, which produces nonlinear electromagnetic coupling effects not fully captured by the aggregate AID and forward power features, as documented by Zhang et al. [7].
Comment 3: In the text, only a single measurement batch is excluded for LOSO, and the physical specimens are still from the same batch of steel plates. Only multiple measurements are repeated, and there is no completely independent new corrosion specimen. The model learns the inherent electromagnetic base characteristics of the same steel plate, rather than pure corrosion characteristics. The ultra-high accuracy may have false optimism and cannot prove that the model can generalise to completely new, unknown corroded steel plates.
How we addressed:
“In the text, only a single measurement batch is excluded for LOSO, and the physical specimens are still from the same batch of steel plates.”
Ans: LOSO validation scope. We fully agree that the LOSO validation operates at the readcount level and does not provide specimen-level independence, since all records originate from the same four physical steel plates. This is now explicitly acknowledged in the revised manuscript. The limitations section states clearly that the model may capture plate-specific electromagnetic baseline characteristics alongside corrosion-induced changes, and that validation on entirely independent steel plates from separate batches is necessary to confirm true generalisation to unknown corroded specimens.
Only multiple measurements are repeated, and there is no completely independent new corrosion specimen.”
Ans: The absence of independent specimens is a constraint inherited from the dataset itself rather than a methodological choice. The dataset is a publicly available benchmark originally published by Zhang et al. and used across several peer-reviewed studies, in which only one physical steel plate per corrosion stage was available by experimental design. As a secondary dataset study, introducing new independent specimens is beyond the scope of this work. Unfortunately, the limited availability of controlled corrosion specimens and the significant experimental cost associated with generating independently corroded samples under controlled marine atmospheric conditions also preclude the collection of additional data within the scope of the present study. Collecting data from multiple independent corroded specimens across varied material batches is identified as a primary direction for future work.
Comment 4: Corrosion variables are single and mixed, lacking controlled variable control experiments. There is no control group set up, no long-term environmental control of blank non-corroded steel plates, and no experiment on the interference of humidity/temperature on RFID signals, making it impossible to distinguish whether signal changes are caused by corrosion of the oxide layer or environmental noise.
How we addressed:
We thank the reviewer for this observation. The absence of controlled variable experiments, blank reference specimens, and humidity/temperature isolation is a limitation of the original experimental design of Zhang et al. [7], from which this publicly available benchmark dataset was sourced. As a secondary dataset study, modifying the underlying experimental conditions is beyond the scope of this work. It is also worth noting that the same International Paint S275 mild steel samples and experimental protocol have been used across multiple peer-reviewed studies in this research area, including related electromagnetic NDT investigations reported in references [9] and [13], where similar environmental conditions without blank reference specimens or humidity/temperature isolation were equally applied. The experimental scope considered in those studies, including the absence of environmental baselines and controlled variable isolation, is addressed in separate publications and is therefore beyond the scope of the present secondary dataset analysis. This limitation is already acknowledged in the limitations section of the manuscript, which explicitly notes that real-world factors, including humidity, temperature variation, and electromagnetic interference, were not included. Designing controlled experiments with environmental baselines and reference specimens is identified as an important direction for future work.
Comment 5: Experimental sweep frequency is 902-928MHz, and corrosion changes the resonant frequency of the antenna. Theoretically, the resonant shift varies with different degrees of corrosion, and the frequency itself should have discriminative ability. However, SHAP's frequency contribution of 0 only indicates the limitations of the fixed frequency scanning process design in this experiment, which does not have universality and cannot be extended to variable frequency and multi-resonant RFID tags. The conclusion is to use a special case instead of a general rule.
How we addressed:
We thank the reviewer for this precise observation and agree entirely. The zero SHAP contribution of frequency is an artefact of the fixed frequency sweep design of the original experimental protocol, in which the same 902--928 MHz sweep was applied identically across all corrosion stages, making frequency non-discriminative by construction rather than by physical principle. Two clarifications have been added to the revised manuscript: a sentence in the SHAP discussion section explicitly stating that this result reflects the fixed sweep design limitation rather than a general physical conclusion, and a sentence in the limitations paragraph noting that this finding should not be generalised to variable frequency or multi-resonant RFID sensing systems, where corrosion-induced resonant frequency shifts would be expected to carry discriminative information. This interpretation is further supported by new text added to the SHAP discussion section (lines 556--562), with a new reference added as reference [23]: Zhang, H., He, Y., Gao, B., Tian, G.Y., Xu, L. and Wu, R., "Evaluation of atmospheric corrosion on coated steel using K-band sweep frequency microwave imaging," IEEE Sensors Journal, vol. 16, no. 9, pp. 3025--3033, 2016. Zhang et al. [23] demonstrated in K-band sweep-frequency microwave NDT that selecting individual frequency points for corrosion characterisation is suboptimal and that the aggregate response across the full frequency sweep carries the dominant diagnostic information.
Comment 6: Laser thickness measurement and RFID RF acquisition both have inherent instrument noise, which theoretically cannot be fully filed; The ultra-high accuracy comes from the fact that the dataset only has 4 discrete thickness labels, which belong to limited classification fitting rather than continuous regression. The model is essentially a classification disguised as regression, and the ability to predict true continuous thickness is overestimated.
How we addressed:
We thank the reviewer for this important observation. We agree that both laser profilometry thickness measurement and RFID RF acquisition contain inherent measurement noise, and that a regression model cannot completely remove these uncertainties. We also agree that the available dataset contains four nominal corrosion thickness labels, namely 0, 43, 77, and 108 µm, corresponding to the M0, M1, M3, and M6 corrosion stages. This distinction has been made explicit from the Literature Review section onward, where the task is described as estimating nominal corrosion thickness states derived from experimentally measured corrosion layers (Section 2, lines 124--125). Therefore, we have revised the manuscript to avoid implying that the model has been validated for arbitrary continuous thickness prediction across an unrestricted corrosion-thickness range. The purpose of the regression analysis in this study is to estimate nominal corrosion layer thickness values associated with experimentally defined corrosion stages, rather than to claim full continuous-thickness generalisation.
The thickness labels used in the present work were not artificially generated; they were taken from direct laser profilometry measurements reported in the foundational experimental study by Zhang and Tian [7], where average corrosion layer thicknesses of 43, 77, and 108 µm were measured for the 1-, 3-, and 6-month exposed samples. The present manuscript now states that the corrosion thickness labels were assigned as 0, 43, 77, and 108 µm based on these prior laser profilometry measurements (Section 3.4, lines 329--332).
We have also revised the terminology throughout the manuscript to refer to "nominal thickness estimation" rather than to “unrestricted continuous thickness prediction”. This clarification makes explicit that the regression model is trained and validated within a small number of experimentally measured nominal thickness states. The revised manuscript in Section 4.2.1 (lines 444--453) now explicitly states that the reported regression performance should be interpreted as accurate estimation of nominal corrosion thickness states within the available label space rather than unrestricted prediction across a continuous corrosion-thickness spectrum, and that validating predictive capability beyond the experimentally observed thickness levels remains an important direction for future research.
The revised limitations section in Section 5 (lines 609--617) further acknowledges that the dataset comprises a small label space of four experimentally measured nominal thickness values derived from a single corrosion mechanism, that the available thickness labels represent discrete nominal corrosion states rather than a densely sampled continuous thickness distribution, and that validation using additional independently measured corrosion thickness levels and independent physical specimens will be required to establish continuous-regression capability and broader generalisation. In addition, we added a discussion of the 43 µm M1 stage, where the prediction uncertainty is highest. The revised manuscript explains that early-stage corrosion at this thickness level is physically more variable because the oxide layer is thin, spatially non-uniform, and affected by variable dielectric properties and surface roughness. This is consistent with Zhang and Tian [7], who showed that corrosion-induced changes in surface roughness, conductivity, and permeability influence RFID antenna impedance and therefore produce nonlinear electromagnetic coupling effects.
Comment 7: The derivation of the AID formula assumes that the reader tag distance and antenna polarisation are completely cancelled out, but in reality, the edges of metal components, multiple layers of rust, and non-uniform oxide layers will introduce additional coupling losses. The paper does not discuss the failure boundary conditions of AID and excessively claims that AID is not affected by geometric posture. Only explaining the positive correlation between forward power, AID, and corrosion thickness, without explaining the characteristic nonlinear fluctuations caused by the dielectric constant, roughness, and uneven corrosion of the oxide layer, cannot explain the source of prediction error for M1 (43 µ m), and the combination of mechanism and model results is incomplete.
How we addressed:
“The derivation of the AID formula assumes that the reader tag distance and antenna polarisation are completely cancelled out, but in reality, the edges of metal components, multiple layers of rust, and non-uniform oxide layers will introduce additional coupling losses.”
Ans: We thank the reviewer for this important observation. We agree that the AID formulation is derived under idealised propagation assumptions and should not be interpreted as providing complete immunity to reader--tag orientation, distance, or environmental effects. The original work of Zhang and Tian [7] introduced AID to reduce the influence of reader--tag orientation, distance, and environmental variability through sweep-frequency measurements and PCA-based feature extraction. Still, it did not claim complete elimination of all measurement influences. To address the reviewer's concern, we have revised the manuscript in three locations.
“The paper does not discuss the failure boundary conditions of AID and excessively claims that AID is not affected by geometric posture.”
Ans: In Section 3.2 (lines 245--252), the manuscript now explicitly states that the geometric compensation provided by AID is only approximate and is derived under idealised line-of-sight propagation assumptions. In practical corrosion-monitoring environments, additional electromagnetic effects including multipath propagation, polarisation mismatch, metallic edge diffraction, chip nonlinearity, multiple rust layers, and non-uniform oxide growth may introduce coupling losses that are not fully cancelled within the AID formulation. Consequently, AID should be interpreted as a feature with reduced sensitivity to reader-tag geometry rather than complete immunity to geometric and environmental influences.
“Only explaining the positive correlation between forward power, AID, and corrosion thickness, without explaining the characteristic nonlinear fluctuations caused by the dielectric constant, roughness, and uneven corrosion of the oxide layer, cannot explain the source of prediction error for M1 (43 µ m)”
Ans: In Section 4.2.1 (lines 456--465), the revised manuscript expands the discussion of the M1 (43 µm) prediction uncertainty, explaining that early-stage corrosion is characterised by thin, spatially heterogeneous oxide layers strongly influenced by local variations in oxide-layer dielectric constant, surface roughness, conductivity, permeability, and corrosion morphology. Localised pitting and non-uniform oxide growth can generate spatially varying impedance distributions across the sensing region, producing nonlinear fluctuations in the RFID response that are not completely represented by the aggregate AID and forward-power features, resulting in greater electromagnetic variability and prediction uncertainty than the more developed corrosion stages. In Section 5 (lines 633--639), the limitations paragraph now states that the AID formulation is derived under idealised propagation assumptions and therefore provides partial rather than absolute compensation of reader-tag orientation and distance effects. Practical factors such as edge diffraction from metallic components, multipath propagation, polarisation mismatch, oxide-layer inhomogeneity, localised pitting corrosion, and multilayer rust growth may introduce additional coupling effects that are not explicitly modelled, and their influence on AID stability and machine-learning performance remains an important topic for future experimental investigation.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript aims to help to reduce the consequences of the issue with metallic corrosion. The authors propose a combination of a well-established method (UHF RFID) for corrosion detection and a modern implementation of the algorithmic modelling, i.e. machine learning.
The topic of the manuscript is relevant, but the actual execution of the idea leaves much to be desired. It is described in the manuscript that steel specimens were allowed to corrode in marine atmosphere for 0, 1, 3, 6, 10 and 12 months. How many replicates were performed to ensure reproducibility? How were specimens prepared? Were they polished? As the authors state, the surface roughness affects corrosion processes, but they did not account for the initial roughness of the specimens. Why were some specimens allowed to corrode for 10 and 12 months if only early corrosion was researched?
The authors cite and reproduce the work of Zhang et. all, reference 7 in the manuscript, and use the data about corrosion products thickness obtained experimentally in the cited work. The authors of reference 7 state that they observed pitting corrosion but ignored, as the authors of the present manuscript, that pitting corrosion is a localized form of corrosion and does not form a layer of uniform thickness of corrosion products, worse – it digs invisibly toward the bulk of the metal. Math might be correct, but it is not appropriate to mix newly obtained experimental data with old outside data, from a replicated in the manuscript research. This mixture makes the results and the following conclusions uncertain. To validate the proposed method, data for the thickness of corrosion products layer, combined with tag reads of the same specimens from enough reliable physical experiments must be obtained.
Some figures include labels a, b,... while others do not.
A formula is missing following the colon on line 217.
Author Response
Comment 1: The topic of the manuscript is relevant, but the actual execution of the idea leaves much to be desired. It is described in the manuscript that steel specimens were allowed to corrode in a marine atmosphere for 0, 1, 3, 6, 10 and 12 months. How many replicates were performed to ensure reproducibility? How were specimens prepared? Were they polished? As the authors state, the surface roughness affects corrosion processes, but they did not account for the initial roughness of the specimens. Why were some specimens allowed to corrode for 10 and 12 months if only early corrosion was researched?
How we addressed:
We thank the reviewer for these observations and we have carefully addressed each points.
How many replicates were performed to ensure reproducibility? How were specimens prepared? Were they polished? As the authors state, the surface roughness affects corrosion processes, but they did not account for the initial roughness of the specimens.
Ans: The present study does not involve any new experimental data collection. The specimens, preparation protocol, corrosion exposure conditions, and all associated experimental decisions are sourced entirely from the original experimental study of Zhang et al. [7], which is a publicly available benchmark dataset used across several peer-reviewed studies. Questions regarding the number of replicates, specimen polishing, and initial surface roughness characterisation therefore relate to the original experimental design of Zhang et al. and are beyond the scope of this secondary dataset analysis. As a secondary dataset study, modifying the underlying experimental conditions or introducing additional controlled variables is not within the scope of the present work.
Why were some specimens allowed to corrode for 10 and 12 months if only early corrosion was researched?
Ans: The M10 and M12 samples were part of the original dataset published by Zhang et al. [7] and were collected as part of their broader corrosion characterisation study covering multiple exposure durations. As stated in Section 3.1 (lines 187--191), the M10 and M12 samples remain part of the original dataset but were intentionally excluded from the present analysis because the focus of this work is early-stage corrosion monitoring. Restricting the analysis to M0, M1, M3, and M6 allows evaluation of whether RFID-derived features can identify and estimate corrosion progression before substantial material degradation occurs. The decision to include M10 and M12 in the original data collection by Zhang et al. was part of their broader experimental design and is not a choice made by the authors of the present study.
Comment 2: The authors cite and reproduce the work of Zhang et al., reference 7 in the manuscript, and use the data about corrosion products thickness obtained experimentally in the cited work. The authors of reference 7 state that they observed pitting corrosion but ignored, as the authors of the present manuscript, that pitting corrosion is a localised form of corrosion and does not form a layer of uniform thickness of corrosion products; worse, it digs invisibly toward the bulk of the metal. Math might be correct, but it is not appropriate to mix newly obtained experimental data with old outside data from replicated research in the manuscript. This mixture makes the results and the following conclusions uncertain. To validate the proposed method, data for the thickness of the corrosion products layer, combined with tag reads of the same specimens from sufficiently reliable physical experiments, must be obtained.
How we addressed:
The authors of reference 7 state that they observed pitting corrosion but ignored, as the authors of the present manuscript, that pitting corrosion is a localized form of corrosion and does not form a layer of uniform thickness of corrosion products; worse, it digs invisibly toward the bulk of the metal.
Ans: We agree with the reviewer's observation that pitting corrosion is localised and does not produce a uniform thickness layer. This has been explicitly acknowledged in the revised manuscript in Section 3.4 (lines 329--332), which now states that the thickness values represent average nominal corrosion-product thicknesses obtained from laser profilometry measurements reported by Zhang et al. [7], and that because early corrosion is characterised by localised pitting and spatially non-uniform corrosion-product growth, these values are interpreted as nominal average corrosion states used for regression modelling rather than uniform thickness measurements. The regression task is therefore framed throughout the revised manuscript as nominal thickness state estimation rather than continuous uniform thickness prediction.
Math might be correct, but it is not appropriate to mix newly obtained experimental data with old outside data from replicated research in the manuscript. This mixture makes the results and the following conclusions uncertain. To validate the proposed method, data for the thickness of the corrosion products layer, combined with tag reads of the same specimens from sufficiently reliable physical experiments, must be obtained.
Ans: We thank the reviewer for raising this concern and would like to respectfully clarify the data sources used in this study. The present work does not involve any newly collected experimental data. Both the RFID tag readings and the corrosion layer thickness values of 43, 77, and 108 µm originate from the same publicly available benchmark dataset published by Zhang et al. [7], in which the thickness values were obtained by laser profilometry on the same physical steel specimens from which the RFID readings were collected, all within the same original experimental study. As a secondary dataset study, our contribution is entirely the feasibility study of the machine learning analysis framework applied to this existing dataset, and no mixing of independently obtained data sources has taken place. We hope this clarification addresses the reviewer's concern, and we acknowledge that this secondary dataset nature of the study could have been stated more clearly in the original manuscript. The revised manuscript now makes this explicit from the outset of Section 3.1, where it is stated at lines 195--203 that the 1,089 entries in the dataset represent RFID measurement records rather than independent physical corrosion specimens, that each corrosion-stage file contains multiple RFID reads collected during frequency-sweep measurements repeated several times for the same underlying specimen, and that multiple records originate from the same physical specimen under comparable experimental conditions, all sourced from the original Zhang et al. [7] dataset.
Comment 3: Some figures include labels a, b,... while others do not
How we addressed:
All figures containing multiple panels (Figures 2 and 5) have been updated with subfigure labels and independent sub captions.
Comment 4: A formula is missing following the colon on line 217.
How we addressed:
The missing equation defining the complex modulus has been added immediately after the colon in Section 3.2 (lines 257--259), completing the sentence that previously ended abruptly. The denominator |Z_A[ψ] + Z_L| is now explicitly defined through Equation (7), which presents the complex modulus as the square root of the sum of the squared real and imaginary parts of the combined impedance, confirming that AID remains real-valued and positive for all corrosion states.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors must clarify the environmental conditions under which the dataset was gathered. Introduce a discussion or a baseline test demonstrating how the classifier performs when an external geometric reflector (such as a moving metallic or dielectric object) is placed near the transmission path. Provide proof of structural robustness against spatial misalignment. The authors should evaluate or explicitly discuss the system's sensitivity boundaries: what happens to the classification accuracy when the reader antenna is tilted. For an NDT application, understanding the underlying electromagnetic physics is crucial. The paper fails to explain why certain features (e.g., specific frequency channels or read rate thresholds) carry higher weights during feature selection. For instance, Random Forest feature importance charts or SHAP (Shapley Additive exPlanations) values would bridge the gap between empirical data-driven classifications and the underlying physical antenna mechanics.
Author Response
Comment 1:
“The authors must clarify the environmental conditions under which the dataset was gathered.”
Ans: The experimental conditions have been clarified in the manuscript. The dataset was collected under controlled laboratory conditions with fixed reader-tag geometry and without intentionally introducing external reflectors or environmental disturbances. This enabled isolation of corrosion-induced RFID variations from environmental effects. We have already mentioned this in the paper in lines 180–181 that measurements were conducted in a typical indoor office environment containing metallic furniture and other sources of multipath clutter [7]. However, we have now explicitly explained this in more detail between lines 182–187.
Addressed in the manuscript:
"All RFID measurements were acquired under controlled laboratory conditions. During data acquisition, the relative positions of the RFID reader, antenna, and test samples were maintained constant to minimise geometric variability. No intentional metallic or dielectric reflectors were introduced into the reader–tag path. Therefore, the dataset primarily captures variations caused by corrosion progression rather than environmental perturbations."
Comment 2:
“Introduce a discussion or a baseline test demonstrating how the classifier performs when an external geometric reflector (such as a moving metallic or dielectric object) is placed near the transmission path.”
Ans: We agree that nearby metallic or dielectric objects can influence RFID measurements through multipath propagation, reflection, and signal scattering. A dedicated baseline experiment using controlled external reflectors would indeed provide valuable quantitative insight into the robustness of classifiers under environmental perturbations. However, the objective of this study was to evaluate corrosion progression under controlled laboratory conditions, where the positions of the RFID reader, antenna, and test samples were kept fixed to minimise geometric variability and isolate corrosion-induced changes. Consequently, environmental reflector effects were intentionally excluded from the experimental design.
To address this limitation, we have now explicitly discussed the potential influence of surrounding objects in Section 4.1 and cited established RFID propagation studies demonstrating how nearby conductive or dielectric objects affect RFID link performance.
.
Addressed in the manuscript (Section 4.1, lines 420–424):
"RFID measurements are inherently influenced by the surrounding electromagnetic environment. Previous RFID propagation studies have shown that nearby conductive or dielectric objects can introduce multipath scattering, reflection, and blockage effects that alter received signal strength and backscatter performance. Consequently, an external reflector placed near the reader–tag path may affect measured RFID features independently of corrosion [26]."
Reference [26]: Griffin, J.D.; Durgin, G.D. Complete link budgets for backscatter-radio and RFID systems. IEEE Antennas and Propagation Magazine 2009, 51, 11–25.
Comment 3:
“Provide proof of structural robustness against spatial misalignment.”
Ans: The manuscript now explicitly discusses the influence of reader-tag alignment and polarisation mismatch on RFID measurements. To isolate corrosion effects, a fixed measurement geometry was maintained throughout dataset acquisition.
Addressed in the manuscript (Section 3.1, Dataset Acquisition, lines 187–191):
"RFID system performance depends on the relative geometry between reader and tag antennas. Variations in alignment can change antenna coupling and polarisation matching, leading to variations in received signal strength and communication reliability. Previous RFID studies have identified polarisation mismatch and geometric orientation as important contributors to RFID link variability [2,7,11]."
Comment 4:
“The authors should evaluate or explicitly discuss the system's sensitivity boundaries: what happens to the classification accuracy when the reader antenna is tilted.”
Ans: We agree that reader antenna orientation may influence RFID measurements through polarisation mismatch effects. This limitation has now been acknowledged in the manuscript. Since antenna orientation remained fixed during all measurements, the reported classification accuracy reflects corrosion discrimination rather than orientation-induced variability.
Addressed in the manuscript (Section 3.1, Dataset Acquisition, lines 191–198):
" Also, antenna tilt can alter the polarisation relationship between the reader and RFID tag. RFID link-budget studies have shown that polarisation mismatch introduces additional signal attenuation and may reduce received power levels. Consequently, significant reader antenna tilt may influence feature values used by the classifier [22]. Since the present experiments were conducted using a fixed antenna orientation and reader and sample positioning were maintained constant throughout data acquisition, therefore, classification results represent corrosion-induced variability under fixed geometric conditions."
Reference [22] (Jouali, R.; Ouahmane, H.; Khan, J.; Liaqat, M.; Bhaij, A.; Ahmad, S.; Haddad, A.; Aoutoul, M. Improved stable read range of the RFID tag using slot apertures and capacitive gaps for outdoor localization applications. Micromachines 2023, 14, 1364) has been newly added to the reference list to support this discussion.
Comment 5:
“For an NDT application, understanding the underlying electromagnetic physics is crucial.”
Ans: We thank the reviewer for this important observation. To address this concern, we have added a discussion explaining the physical basis of RFID corrosion sensing. Specifically, corrosion alters the electromagnetic properties of the steel substrate, including conductivity, magnetic permeability, surface roughness, and corrosion-product morphology. These changes affect the interaction between the RFID tag antenna and the metallic surface, leading to variations in impedance matching, coupling efficiency, and backscatter response. Consequently, the RFID-derived features used by the machine learning model are linked to corrosion-induced electromagnetic changes rather than representing purely empirical indicators. We have incorporated this explanation in Section 4.1 and cited prior RFID corrosion studies that demonstrate the influence of conductivity and permeability changes on RFID sensing performance, including two newly added references.
Addressed in the manuscript (Section 4.1, lines 411–418):
"The physical basis of RFID corrosion sensing arises from the interaction between the tag antenna electromagnetic field and the corroding metallic substrate. As corrosion develops, the local conductivity, magnetic permeability, surface roughness, and corrosion-product morphology of the steel surface change [7,10,24]. These changes modify the antenna boundary conditions, impedance matching, coupling efficiency, and backscatter response of the RFID tag [25]. Therefore, variations in RFID-derived features are not purely empirical but originate from corrosion-induced changes in the electromagnetic interaction between the tag antenna and the steel substrate. Previous LF RFID corrosion studies have similarly shown that conductivity and permeability variations in corrosion layers influence RFID impedance matching and sensing performance [10,24]."
Two new references were added to support this discussion: reference [24] (Sunny, G.A. Passive low frequency RFID for non-destructive evaluation and monitoring. PhD thesis, Newcastle University, 2017) and reference [25] (Soodmand, S.; Zhao, A.; Tian, G.Y. UHF RFID system for wirelessly detection of corrosion based on resonance frequency shift in forward interrogation power. IET Microwaves, Antennas & Propagation 2018, 12, 1877–1884).
Comment 6:
“The paper fails to explain why certain features (e.g., specific frequency channels or read rate thresholds) carry higher weights during feature selection.”
Ans: We thank the reviewer for this valuable comment. To improve the physical interpretability of the machine learning model, we have added a discussion explaining the significance of the most influential RFID features. In particular, we clarify that forward power (threshold power) is physically meaningful because it represents the minimum reader power required to activate the RFID tag. Corrosion-induced changes in conductivity, magnetic permeability, and surface morphology can modify the antenna impedance and coupling conditions between the tag and the metallic substrate. These changes affect power-transfer efficiency and consequently alter the activation power required for successful tag operation. Therefore, the high importance assigned to forward power is consistent with established RFID link-budget theory, where impedance mismatch and coupling efficiency directly influence delivered tag power and read performance. This explanation has been added to the manuscript to provide a clearer connection between feature importance and the underlying sensing physics.
Addressed in the manuscript (Section 4.3, lines 565–572):
"Forward power is physically meaningful because it represents the reader power required to activate the RFID tag. Corrosion-induced changes in conductivity, permeability, and surface morphology modify the tag antenna impedance and coupling conditions. This alters the power-transfer efficiency between the reader and tag, meaning that more severely corroded samples may require higher activation power. Therefore, the high importance assigned to forward power is consistent with RFID link-budget theory [26], where impedance mismatch and antenna coupling directly affect delivered tag power and read performance."
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsAccept
Author Response
Thank you
Reviewer 3 Report
Comments and Suggestions for AuthorsThe authors have eliminated the drawbacks from the previous version and now the manuscript has a solid sintific foundation, so I recomend to publish it.
Author Response
Thank you
